INQUIRING LINE

Nearly every LLM benchmark in a 445-paper review fell short on measuring what its name promises: how should researchers document those gaps?

How should researchers document validity gaps in their benchmarks?

This explores what researchers should write down about the limits of their benchmarks, meaning the places where a score might not measure what it claims to, so that readers can judge how far to trust it.


This explores what researchers should write down about the limits of their benchmarks, so a reader knows how much weight a score can bear. The corpus has no ready-made documentation template. It does have something more useful: a map of where the gaps usually hide, which gives you a checklist. The baseline is grim. A review of 445 LLM benchmark papers found that almost all of them fall short on construct validity, the question of whether a test measures the thing it's named after Do LLM benchmarks actually measure what they claim to measure?. The recurring problems point straight at what to document. Define the ability you're testing. Explain why a task borrowed from older work still fits. Report statistical tests. Say plainly which claims your scores support and which they don't.

The least obvious gap is what you threw away. Benchmarks routinely drop examples where human annotators disagreed, which seems like good data hygiene. But that filtering removed exactly the cases that expose LLM weakness at handling ambiguity. On ambiguous examples, accuracy was 32%, against 90% on the cleaned versions, and standard evaluation never shows the difference Do standard NLP benchmarks hide LLM ambiguity failures?. So a validity statement should list the curation choices and what they excluded, not just what's left. The same goes for scoring choices. LR²Bench scores only final answers against fixed correct answers, because grading the reasoning steps would reward models that imitate the style of reasoning without getting anything right Should reasoning benchmarks score final answers or reasoning traces?. Each choice like that is a validity decision, and it should be written down with its reason.

Contamination is the gap readers care about most and authors document least. One model could rebuild 54.6% of the MATH-500 problems from partial prompts, yet scored 0% on a newer benchmark released after its training data was collected Does RLVR success on math benchmarks reflect genuine reasoning improvement?. A companion note adds a twist: a training method can genuinely change how a model reasons while the benchmark gains still come from memorization. The two happen at different levels of measurement Can genuine reasoning activation coexist with contaminated benchmarks?. A good report says which one its scores can speak to. Running the same test on data released after the model's training cutoff is the cheapest check available.

A second idea in the corpus is that documentation shouldn't be prose alone. It can travel with the score as evidence. Benchmark Radar keeps each score linked to its source and original setup, because a number with no record of where it came from can't support fair comparisons between models Can benchmark scores be trusted without knowing their origin?. BenchShield goes further for agent benchmarks. Benchmark operators can claim an agent "validly completed the task" based on recorded logs of how it ran, not just a final pass/fail number Can infrastructure evidence replace terminal scores in benchmark validation?. That matters because switching to interactive, multi-step evaluation doesn't remove these problems. Comparability, reproducibility, and how evidence maps to a verdict all show up again across the whole sequence of an agent's actions, and they need shared protocols to stay interpretable Do interactive evaluations actually solve the benchmark comparison problem?.

The idea worth borrowing comes from AI paper-writing systems. Spark-to-Paper requires the system to specify what evidence would count before it sees any results Can separating judgment from verification improve research paper reliability?. That is essentially preregistration, the practice of committing to an analysis plan in advance. Applied to benchmarks, it would mean writing down before any testing what the benchmark is meant to measure, what result would count as failure, and what the data excludes. That's the honest version of a validity section. Without it, limitations tend to get written after the scores are in, to fit whatever the scores turned out to be.


Sources 9 notes

Do LLM benchmarks actually measure what they claim to measure?

A review of 445 benchmark papers found almost all have gaps in construct validity. Phenomena are often poorly defined, tasks are reused without adjustment, statistical testing is rare, and claims frequently don't follow from scores.

Do standard NLP benchmarks hide LLM ambiguity failures?

By filtering out examples where annotators disagree, benchmarks remove test cases that would reveal LLM failures at ambiguity recognition. Research using ambiguous examples shows a 32% vs. 90% accuracy gap invisible to standard evaluation.

Should reasoning benchmarks score final answers or reasoning traces?

LR²Bench scores only final answers against deterministic ground truth, not reasoning steps. This methodological choice reveals a 20% ceiling that trace-based evaluation would inflate by counting stylistic reasoning mimicry as actual reasoning capability.

Does RLVR success on math benchmarks reflect genuine reasoning improvement?

Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.

Can genuine reasoning activation coexist with contaminated benchmarks?

RLVR activates genuine reasoning patterns through RL training while benchmark improvements may reflect data memorization on contaminated datasets. These operate at different measurement levels and can coexist without contradiction.

Show all 9 sources
Can benchmark scores be trusted without knowing their origin?

Benchmark Radar catalogs AI evaluations while keeping source identities and citations attached, enabling readers to trace scores back to their original settings. The system demonstrates that scores without provenance cannot reliably support model comparisons.

Can infrastructure evidence replace terminal scores in benchmark validation?

BenchShield enables benchmark operators to issue claims about valid task completion grounded in recorded infrastructure evidence rather than terminal scores alone. This shifts from a single number to a verifiable claim about whether an agent followed the intended evaluation path.

Do interactive evaluations actually solve the benchmark comparison problem?

Interactive evaluation relocates core problems—comparability, reproducibility, evidence-to-judgment mapping—into higher-dimensional space rather than solving them. The field needs design protocols and shared standards, not format adoption, to make trajectory scoring interpretable.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.