Why does bragging about a system's best-ever score, out of many tries, make progress look bigger than it really is?
Why do cumulative-best reporting inflate progress compared to actual validation?
This explores why reporting the best score a system has ever reached across many runs, checkpoints, or attempts can make progress look bigger than a fresh, independent check would show. The corpus has no paper that studies cumulative-best reporting by name, but several notes explain how it inflates results.
This explores why reporting the best score a system has ever reached across many runs, checkpoints, or attempts can make progress look bigger than a fresh, independent check would show. The corpus has no paper that studies cumulative-best reporting by name. Several notes still explain why it inflates results, and the shared lesson is that choosing the best result is itself a form of optimization. Optimizing against a weak signal produces gains that may not be real.
The clearest explanation comes from work on reward hacking. It finds that reward hacking doesn't only happen when a model's weights are trained. It also happens when outputs are *selected* and when prompts are revised, and the cause is the same each time: optimizing against a score that only partly captures the real task Does reward hacking always stem from the same failure?. Keeping a running best is output selection over time. Each new attempt is another draw, and keeping the maximum picks up lucky draws along with real improvement. A related note points out that even a 'deterministic' run at temperature zero is still just one sample from the model's range of possible outputs Does setting temperature to zero actually make LLM outputs reliable?. If one run is a single draw, then the best of many draws tells you more about the luck of the draws than about how the system will usually perform.
A second source of inflation is what the score is measured on. One study found that a math model could reconstruct more than half of a popular benchmark from partial prompts, yet scored zero on a benchmark released after its training data was collected Does RLVR success on math benchmarks reflect genuine reasoning improvement?. A companion note argues that genuine reasoning improvement and benchmark gains can happen separately, so a rising score doesn't prove the underlying ability grew Can genuine reasoning activation coexist with contaminated benchmarks?. Put that together with best-of-many selection and you are picking the peak of a signal that may already be inflated by memorization.
The fix that keeps coming up is to validate against something the selection process never saw. A small model trained on whether its harness patches actually worked beat prompted frontier models, because it reran each patch to confirm the effect, while the prompted models optimized for patches that merely looked plausible Does training editors on real outcomes beat prompting larger models?. Spark-to-Paper takes the same idea into research writing: it requires the evidence standard to be set *before* results are seen, so nobody can quietly choose the flattering run afterward Can separating judgment from verification improve research paper reliability?. Committing to the check before seeing results is what separates validation from keeping the best score.
The less obvious point is that more detailed evaluation doesn't fix the problem on its own. Moving from single scores to scoring the agent's full sequence of steps brings the old comparability and reproducibility problems back in a more complicated form Do interactive evaluations actually solve the benchmark comparison problem?. Two tools aim at provenance instead. Benchmark Radar keeps each score's source attached so readers can see what setting produced it Can benchmark scores be trusted without knowing their origin?. BenchShield backs claims about valid task completion with recorded infrastructure evidence instead of a final number alone Can infrastructure evidence replace terminal scores in benchmark validation?. A best-ever number without its history can't be checked. The cure is not a better number but knowing how many attempts went into the number you're shown.
Sources 9 notes
Reward hacking arises during weight training, output selection, and prompt revision through a shared failure: optimization against signals that incompletely represent the actual task. The substrate matters less than the misalignment between the scoring function and ground truth.
Fixed seeds and zero temperature replicate the same output repeatedly, but that output remains one draw from the model's probability distribution. McDonald's omega testing across 100 repetitions reveals that consistency does not equal reliability.
Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.
RLVR activates genuine reasoning patterns through RL training while benchmark improvements may reflect data memorization on contaminated datasets. These operate at different measurement levels and can coexist without contradiction.
A 9B model trained with reinforcement learning on patch success raised a frozen agent's performance by 9.3 points across three tasks, while prompted frontier models produced unstable or lower gains. The difference stems from feedback: trained editors rerun patches to verify impact, while prompted models optimize only for plausibility.
Show all 9 sources
Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.
Interactive evaluation relocates core problems—comparability, reproducibility, evidence-to-judgment mapping—into higher-dimensional space rather than solving them. The field needs design protocols and shared standards, not format adoption, to make trajectory scoring interpretable.
Benchmark Radar catalogs AI evaluations while keeping source identities and citations attached, enabling readers to trace scores back to their original settings. The system demonstrates that scores without provenance cannot reliably support model comparisons.
BenchShield enables benchmark operators to issue claims about valid task completion grounded in recorded infrastructure evidence rather than terminal scores alone. This shifts from a single number to a verifiable claim about whether an agent followed the intended evaluation path.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Spurious Rewards: Rethinking Training Signals in RLVR
- The Invisible Leash: Why RLVR May Not Escape Its Origin
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Interactive Evaluation Requires a Design Science
- AI for Auto-Research: Roadmap & User Guide
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?