Skip to content

LightEval Benchmarks

Complete list of supported benchmarks in the LightEval adapter. Benchmarks can be referenced individually by ID or via named groups (e.g. commonsense_reasoning runs HellaSwag + WinoGrande + OpenBookQA together).

Benchmark IDIncluded tasksDescription
commonsense_reasoningHellaSwag, WinoGrande, OpenBookQACommonsense reasoning aggregate
scientific_reasoningARC Easy, ARC ChallengeScientific reasoning aggregate
physical_commonsensePIQAPhysical commonsense reasoning
truthfulnessTruthfulQA MC, TruthfulQA GenTruthfulness evaluation
mathGSM8K, MATH Algebra, MATH Counting & ProbabilityMathematical reasoning aggregate
knowledgeMMLU, TriviaQAKnowledge evaluation aggregate
language_understandingGLUE CoLA, GLUE SST-2, GLUE MRPCLanguage understanding aggregate
Benchmark IDNameDescription
hellaswagHellaSwagCommonsense NLI sentence continuations
winograndeWinoGrandeWinograd schema challenge
openbookqaOpenBookQAOpen-book science question answering
Benchmark IDNameDescription
arc:easyARC EasyAI2 Reasoning Challenge — easy split
arc:challengeARC ChallengeAI2 Reasoning Challenge — challenge split
piqaPIQAPhysical Intuition QA
Benchmark IDNameDescription
truthfulqa:mcTruthfulQA (MC)Multiple-choice truthfulness evaluation
truthfulqa:genTruthfulQA (Gen)Generative truthfulness evaluation
Benchmark IDNameDescription
gsm8kGSM8KGrade school math word problems
math:algebraMATH — AlgebraCompetition mathematics (algebra subset)
math:counting_and_probabilityMATH — Counting & ProbabilityCompetition mathematics (combinatorics subset)
math_500MATH-500500-problem subset of MATH benchmark
aime24AIME 2024American Invitational Mathematics Examination 2024
aime25AIME 2025American Invitational Mathematics Examination 2025
Benchmark IDNameDescription
mmluMMLUMassive Multitask Language Understanding (57 subjects)
triviaqaTriviaQATrivia question answering
gpqa:diamondGPQA DiamondGraduate-level Google-proof science questions — Diamond (hardest) split; HuggingFace gated dataset (Idavidrein/gpqa)
Benchmark IDNameDescription
glue:colaGLUE CoLACorpus of Linguistic Acceptability
glue:sst2GLUE SST-2Stanford Sentiment Treebank (binary)
glue:mrpcGLUE MRPCMicrosoft Research Paraphrase Corpus
Benchmark IDNameDescription
lcb:codegeneration_v6LiveCodeBench — Code Generation v6Post-cutoff competitive programming problems (contamination-free); metric: codegen_pass@1

For complete documentation, see the LightEval README.