Do LLM benchmarks actually measure what they claim to measure?
A systematic review examined whether 445 LLM benchmarks have sound construct validity—whether their tasks and metrics truly capture the phenomena they're designed to test. This matters because flawed benchmarks can mask model failures and mislead research.
A team of 29 expert reviewers coded 445 LLM benchmark articles from ICML, ICLR and NeurIPS (2018 to 2024) and from ACL, NAACL and EMNLP (2020 to 2024) against a codebook for construct validity, the degree to which "a benchmark score provides evidence for making claims about the target phenomenon." The review finds that "almost all articles have weaknesses in at least one area across phenomena, tasks, metrics, and claims." Its Discussion locates the gaps: "the operationalisation of abstract phenomena was often insufficient, with definitions being missing or contested"; tasks were "frequently taken from pre-existing data sources without adjustments"; "statistical testing was also rarely performed"; and "about half of the reviewed articles did discuss the validity of their benchmark, but nearly every paper had weaknesses in at least one area." These are the review's findings, and the library records them as its findings.
The reasoning runs through a simple decomposition. A benchmark is a task and a metric used together to represent a phenomenon, read alongside the authors' interpretation of the result. The excerpt's example is GSM8K: the phenomenon is "multi-step mathematical reasoning," measured by exact match on grade-school word problems. Construct validity is borrowed from psychological testing and split into face, content, ecological, predictive, and convergent and discriminant validity. The weak points the review finds sit at the joins: whether the phenomenon is defined, whether the task represents it, whether the metric scores it, and whether the claim follows from the score. Its eight recommendations and operational checklist are aimed at those joins, to be used during design and again during interpretation.
Against the library, the nearest note is Do standard NLP benchmarks hide LLM ambiguity failures?, which argues that filtering out ambiguous items makes one LLM failure invisible. This review supplies the wider frame: the excluded items are one instance of a general gap between what a benchmark says it measures and what its tasks and metrics can show. The review's concern with tasks reused from existing data is the same mechanism at a different step. The excerpt does not discuss ambiguity, so it neither confirms nor extends that note's specific claim. The FLASK note, Do all AI skills improve equally as models scale?, describes a score that rises without the underlying skill improving, which is the failure the review's claim-level checks are meant to catch. The Can benchmark scores be trusted without knowing their origin? note places inspectability in a catalog; the review places it in reporting, asking authors to append the checklist with answers and explanations for skipped items. The psychometric note, Do LLMs show reproducible psychological profiles when given standardized tests?, applies the instrument tradition from which this review takes its central concept, so it inherits the same question of whether a score measures what its label claims.
What the excerpt does not establish is also important. The sample is limited to leading peer-reviewed venues, so it excludes industry-lab benchmarks released without review and specialized venues, as the Limitations section says. Screening was partly automated: GPT-4o mini removed articles, validated at an F1 of 84% on 50 articles, and the authors concede it "may have introduced undetected false negative systematic errors." Coding reliability was moderate, with a mean Brennan–Prediger Kappa of .524 across 30 categorical questions, measured on a random sample of 46 papers reviewed twice; most papers had one primary reviewer. The excerpt also does not show that any single benchmark is invalid, or how far a given gap changes the conclusions drawn from it. The implication at the strength the evidence allows: the 445-article count is a strong signal about how benchmarks in these venues are designed and reported, and the checklist is the authors' proposal, not a validated instrument. Whether following it raises validity is untested here.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What gaps exist between benchmark performance and real deployment outcomes?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do standard NLP benchmarks hide LLM ambiguity failures?
When benchmark creators filter out ambiguous examples before testing, do they accidentally make it impossible to measure whether language models can actually handle ambiguity the way humans do?
ambiguity filtering is one instance of the broader construct validity gap this review measures
-
Do all AI skills improve equally as models scale?
Different evaluation skills show strikingly different scaling patterns. Understanding where skills saturate has immediate implications for model deployment and capability requirements across domains.
a score rising for the wrong reason, the failure claim-level checks target
-
Can benchmark scores be trusted without knowing their origin?
When AI researchers report benchmark scores, how much do the evaluation settings, prompts, and data splits affect the results? Why does tracing each score back to its source matter for fair model comparison?
both make benchmark settings inspectable, through a catalog versus a reporting checklist
-
Do LLMs show reproducible psychological profiles when given standardized tests?
When administered psychological instruments repeatedly, do large language models produce consistent, model-specific response patterns? Understanding whether LLMs exhibit stable behavioral signatures matters for characterizing their deployed behavior and detecting systematic differences across models.
applies the psychological-testing tradition that this review takes its construct validity concept from
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Can AI Do Strategy?
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- Can Large Language Models Reason and Plan?
- How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Benchmarking the Pedagogical Knowledge of Large Language Models
Original note title
nearly every one of 445 reviewed LLM benchmarks has a construct validity weakness — phenomena are often poorly defined and tasks often reused unadjusted