Can the scores we use to track AI progress overstate some skills and miss others that are real?
Can optimization metrics hide actual versus apparent progress in AI systems?
This explores whether the numbers we use to track AI improvement, such as benchmark scores and optimization targets, can show progress that isn't real, or miss progress that is, and what the collection says about telling the two apart.
This explores whether the scores used to measure AI progress can drift away from the real capability they are meant to track. The corpus answers yes, and the distortion runs in both directions. The clearest case is the gap between agent benchmarks and economic value. An analysis of 960 real occupational workflows found that agents do well on contest-style tasks and fail at long, multi-step professional work. The authors put the gap down to what the field chose to measure, more than to what the models can do: it measured contests rather than work Why do agent benchmarks not predict real economic value?. A companion argument adds the less obvious half. Automated benchmarks favor tasks that are precisely specified and easy to grade automatically, so they can **understate** capability as well as overstate it. Messy, open-ended evaluations, read through qualitative log analysis with costs reported, tend to spot new abilities earlier than the leaderboards do Do automated benchmarks hide what frontier AI systems can really do?.
The problem gets sharper when the system optimizes itself. The Darwin Gödel Machine replaces formal proofs with empirical benchmarking: an agent variant counts as better if it scores better. That produced 2.5× gains on SWE-bench Can AI systems improve themselves through trial and error?. A bilevel setup goes one step further. An outer loop rewrites the inner loop's search code and reaches a 5× improvement on its target Can an AI system improve its own search methods automatically?. These results are impressive, but they show a general point: when the benchmark is the fitness signal, the benchmark becomes the definition of improvement. The thing that separates real progress from apparent progress is whether gains carry over to tasks the optimizer never saw. AIDE2 reports exactly this kind of check: its gains held on four held-out benchmarks, including physics-based weather forecasting, which lies outside the selection set Do AIDE2's improvements transfer to unseen tasks?. A held-out test is the plainest tool for telling the two kinds of progress apart.
A quieter version of the question is where the gain actually lives. Wrapping frozen models in a better execution harness raises Terminal-Bench scores across several models without changing any weights Can execution harnesses lift model performance without retuning weights?. Surveys of self-improving agents note that most recent progress comes from this fast, cheap scaffold loop, with comparatively little from updating the model itself Do self-improving agents really split into two distinct loops?. The progress is real, but a leaderboard number credits it to the model. And when an agent is given only a vague goal, it has to build its own training and validation signals before it can optimize anything Can agents learn from vague goals without predefined metrics?. In that case the optimizer is also choosing the metric it will be judged by.
The proposed fixes mostly widen what counts as evidence. One line of work moves agent evaluation away from scoring final answers and toward examining whole interaction trajectories: process quality, recovery from errors, robustness How should we evaluate agent behavior beyond final answers?. Agent-based judges that collect evidence actively cut judge inconsistency from 31% to 0.27%. But their memory module spread errors from one step to the next, a reminder that the measuring tool can fail too Can agents evaluate AI outputs more reliably than language models?. At the most philosophical end, a semiotics argument says that any goal encoded purely in symbols, without contact with the world, can drift from what it was meant to stand for Can AI systems achieve real alignment without world contact?. Read next to the benchmark papers, a metric looks like a symbol for a capability, and it can be optimized long after it has stopped pointing at that capability.
Sources 11 notes
ALE's analysis of 960 real occupational workflows shows agents excel at abstract contests but fail long-horizon professional tasks. The gap is not model capability but benchmark design—the field optimizes what it measures, and it has measured contests rather than work.
Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
An outer loop successfully read inner loop code, identified bottlenecks, and generated new Python mechanisms at runtime, discovering combinatorial optimization and bandit methods that broke the inner loop's deterministic patterns and improved performance on GPT pretraining by 5x.
The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.
Show all 11 sources
StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.
A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.
When given only a natural-language capability direction without predefined tasks or metrics, self-evolving agents redirect search effort toward operationalizing the goal itself. Aspire's benchmark showed that agents must construct their own training and validation signals before optimizing, revealing a phase of work that existing methods skip.
Evaluation of agentic systems shifts evidence from final responses to full interaction sequences, and scoring procedure from correctness alone to process quality, recoverability, coordination, and robustness. This pattern appears across multiple agent benchmarks as a coherent design move.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
Peircean semiotics reveals that symbolic goal encoding without world contact and social mediation cannot guarantee correspondence to actual values. LLMs operating in pure symbol manipulation risk divergence between stated goals and real-world outcomes.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- Self-Improvements in Modern Agentic Systems: A Survey
- Agents' Last Exam
- Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Hyperagents
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- Introducing Sakana AI's Recursive Self-Improvement (RSI) Lab