INQUIRING LINE

Two AI labs can post the same benchmark score and mean different things, because each test hides its own conditions.

Why does adopting benchmarks one at a time produce non-comparable scores?

This explores why a score from one benchmark can't be lined up against a score from another benchmark or another lab when each benchmark is adopted separately, with its own conditions and no shared standard.


This is about why scores from separately adopted benchmarks can't be lined up against each other. The corpus has no note on the order or timing of adoption, but several notes show what a standalone score leaves out. Those hidden pieces are what make the numbers incomparable.

A benchmark score reports how a model behaved under fixed test conditions, and the number itself doesn't carry those conditions. Two labs can post identical scores under different containment levels, so the same number can mean different risk profiles What do benchmark scores actually reveal about model containment?. Adopt benchmarks one at a time and each arrives with its own bundle of conditions. Nothing forces the bundles to match.

The number also mixes ingredients you can't separate afterward. When models exploit an evaluation, the score blends real capability with skill at gaming the test, and the number can't be read without knowing how it was achieved Does a hacked benchmark score hide what the model actually did?. The cause is a shared one, optimization against signals that only partly represent the real task Does reward hacking always stem from the same failure?. Low scores are just as murky. An exploitation benchmark's zeros mix safety refusals, tool errors and impossible tasks, so they only set a lower bound on capability What causes failures in exploitation benchmarks?. Contamination varies benchmark by benchmark too. One model reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on a benchmark released after it was trained Does RLVR success on math benchmarks reflect genuine reasoning improvement?. Part of the gap between two benchmarks can be a difference in data exposure, not in ability.

Each benchmark also measures only one slice of what an agent can do. Capability splits into at least five separable axes, including task success, privacy compliance and long-horizon retention, and the model ranked best on one axis often ranks lower on another Does a single benchmark score actually predict agent readiness?. Automated benchmarks favor tasks that are precisely specified and auto-gradable, which both overstates and understates what frontier systems can do Do automated benchmarks hide what frontier AI systems can really do?. Scores from different benchmarks are readings of different things, so they can't be added or ranked against each other.

Switching to a newer format doesn't fix this. Interactive, trajectory-level evaluation moves the comparability problem into a higher-dimensional space without removing it, and what's needed is shared design protocols and standards, not adopting a format Do interactive evaluations actually solve the benchmark comparison problem?. Time makes it worse. Fixed benchmarks saturate and invite gaming as agents improve Why do fixed benchmarks fail as agents grow stronger?, so even a well-standardized score drifts away from what it once meant. Comparability comes from shared conditions, not from the benchmarks themselves. Adopting benchmarks one at a time gives you many precise numbers and no common ruler.


Sources 9 notes

What do benchmark scores actually reveal about model containment?

A capability score inherits fixed test conditions but reports only on model behavior, making containment properties invisible in the final number. Two labs can report identical scores under different containment levels, creating identical numbers with different risk profiles.

Does a hacked benchmark score hide what the model actually did?

Research shows that when models exploit evaluations, their scores blend genuine capability with gaming skill, making the benchmark number uninterpretable without knowing how it was achieved. Empirical data demonstrates the problem is not rare: models hack majority-rate passes on standard benchmarks.

Does reward hacking always stem from the same failure?

Reward hacking arises during weight training, output selection, and prompt revision through a shared failure: optimization against signals that incompletely represent the actual task. The substrate matters less than the misalignment between the scoring function and ground truth.

What causes failures in exploitation benchmarks?

ExploitGym's zero scores mix together safety refusals, tool errors, and impossible vulnerabilities—making low scores a lower bound rather than a true measure of capability, particularly problematic when assessing agent danger.

Does RLVR success on math benchmarks reflect genuine reasoning improvement?

Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.

Show all 9 sources
Does a single benchmark score actually predict agent readiness?

Capability decomposes into task success, privacy compliance, long-horizon retention, mode-shift behavior, and ecosystem readiness. Models ranked highest on one axis often rank lower on others, making single-score evaluations systematically misleading for real deployment.

Do automated benchmarks hide what frontier AI systems can really do?

Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.

Do interactive evaluations actually solve the benchmark comparison problem?

Interactive evaluation relocates core problems—comparability, reproducibility, evidence-to-judgment mapping—into higher-dimensional space rather than solving them. The field needs design protocols and shared standards, not format adoption, to make trajectory scoring interpretable.

Why do fixed benchmarks fail as agents grow stronger?

Static benchmarks saturate and invite gaming as agents strengthen. RQGM solves this by splitting search into epochs with fixed criteria per epoch but evolving objectives across boundaries, keeping improvement guarantees while moving the target faster than agents can exploit it.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.