Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT
Source: Cunningham, Shetty, Cheng, Rush, METR · 2026-07-21
We propose a measure of an AI agent’s optimization ability with an “expenditure horizon.” We give an empirical illustration from the NanoGPT speedrun.
A critical question is the degree to which AI will accelerate its own progress.
Over the last year there have been many credible claims that AI has begun accelerating algorithmic progress but it is hard to tell by how much, and what to expect in the future.1 We have many imperfect sources of evidence on AI’s acceleration of AI R&D:
AI R&D benchmarks. There are a few high-quality benchmarks which test agents on their ability to optimize AI training algorithms. However most do not report any human baseline, or only a baseline score at 8 hours or 40 hours. Additionally, the benchmark problems are often not reflective of frontier-level AI R&D. It seems plausible that agents can get good solutions to many “textbook” or “toy” problems without being able to contribute to highly-optimized frontier algorithms.
Contributions to frontier optimization problems. There have been many recent reports of autonomous or AI-assisted contributions to frontier optimization problems (some of which are AI R&D problems) e.g. TTT-Discover, AlphaEvolve, and LLM-assisted NanoGPT speedrun contributions. However it is difficult to assess the magnitude of these contributions relative to human effort, and reporting is biased towards successes making the results hard to generalize.
Acceleration in capabilities progress. The evidence above was regarding AI’s contribution to AI R&D inputs; we can additionally look at the trajectory of output, e.g. looking at acceleration in capabilities as measured by METR’s time-horizon or Epoch’s ECI. Estimating how much progress is attributable to AI R&D requires also estimating the contribution of human R&D and training compute (e.g. as done in the Mythos Preview system card). This is intrinsically a lagging metric, only observed after the capabilities are realized in a model.
The method in this note is a novel metric for summarizing an agent’s ability on a optimization problem. It thus can be used to summarize performance on existing AI R&D benchmarks (e.g. RE-bench, MLE-bench), but we also show how to apply it to agentic contribution to certain frontier optimization problems, and we illustrate with NanoGPT.
An agent’s “expenditure horizon” gives a quantitative measure of agentic optimization ability.
We define an agent’s “expenditure horizon” over an optimization problem as the dollar value at which the improvement to the goal metric is equal to the improvement by a human with the same budget.
We give a proof-of-concept showing how to measure expenditure horizon on the NanoGPT speedrun, by comparing human and agentic scaling curves.
We conducted two interviews with prolific NanoGPT contributors. Their responses imply the effort required for an incremental 1 percentage-point optimization is around 16 hours of labor, or $2,400 (at $150/hour). We also use an LLM judge to categorize recent contributions to NanoGPT, which estimates a roughly similar return.
We ran six high-expenditure agentic optimization runs starting from record #78 of the speedrun (March ’26, 85.56 seconds of training time). On reproduction our baseline starts slightly slower which is typically expected on NanoGPT due to noise and hardware differences (Appendix D).
The full returns-to-expenditure curves are shown below, and for each we illustrate the expenditure horizon, shown as the intersection with the estimated returns to human expenditure curve ($2,500/1%).
Our harness is likely inefficient. We ran the agents with continuous access to 4 H100 nodes, to use for validation of their optimizations. As a consequence the agents ran many experiments, likely inefficiently, and experiment cost comprised around 70-90% of the cost of most trajectories. We expect a more optimized harness would significantly lower cost for a given optimization. However we can see from the curves that shifting the expenditure curves horizontally would not dramatically change the expenditure horizon, and it appears unlikely to increase the maximum speedup achieved.
The raw trajectories overstate progress. The trajectories show the cumulative best score. But for each run we also re-validate their trajectories. The overstatement is due to statistical noise. For the best performing models we further think only ~70% of contributions are mergeable as per the maintainer’s judgment.
GPT-5 and Opus-4.1 only chase noise. The raw trajectories from GPT-5 and Opus-4.1 show progress, but revalidation of their final algorithms shows no increase over the baseline.
Some models show significant improvements. GPT-5.5 and Opus-4.8 show significant improvements over the baseline, and we can quantify those with significant expenditure horizons, as shown above.
The big-picture improvements are still modest. Although these models have expenditure horizons in the thousands of dollars, they are small relative to the overall expenditure on human labor. Thus this evidence implies that autonomous optimization does not have dramatic effects on AI R&D progress on NanoGPT (though it is still possible that agents could dramatically augment human progress, AKA hybrid optimization).
Since its launch in May 2024, the leaderboard has compressed training time from approximately 45 minutes to under 2 minutes. The following plot shows all 82 NanoGPT records from May 2024 to April 2026. In an earlier note we estimated how non-obvious each contribution was (depth) and where ideas originated (provenance). We found that humans achieved a 33x speedup over 2024-2026 with deep ideas appearing throughout. Early contributions imported/adapted ideas to catch up to the frontier, while later ones were increasingly invented.
Overall progress during this time was a 57% reduction in training time made up of many small contributions, none contributing more than 8%.
Although the slope of individual contributions varies significantly, the average slope over any significant period is remarkably stable. We estimate this to be around 16 hours per 1% improvement.
The (inverse) semi-elasticity is thus $2,500 of human expenditure per 1% improvement, which we treat as our rough estimate of the local rate of human progress to benchmark agents against. Given the difficulty of estimating human time invested, we note that this is highly uncertain.
The curves above show Opus-4.1 and GPT-5 find no or few optimizations at a low cost, and then plateau. Newer models (GPT-5.2, GPT-5.5, Opus-4.8) continue to increase in a roughly log-linear manner with expenditure into the thousands of dollars.
We estimate expenditure horizons between $0 and $3,300 for models on the NanoGPT speedrun.
Here, we estimate a model’s expenditure horizon as the expenditure needed to achieve the same speedup as the human, at the fitted rate of ~$2,500 per 1% improvement — the intersection between the agent and the human line. For the two frontier models, a star marks the estimated mergeable share of the re-validated win, based on the maintainer’s review of their contributions (discussed below).
Agent progress typically follows an L-shaped curve as suggested by our earlier post on an apple-picking model of AI R&D. However agents seem quite inefficient compared to humans. We think this is a consequence of our harness being inefficient. Agents ran many experiments accounting for ~70-90% of total cost, likely inefficiently, since they were given continuous access to 4 H100 nodes. We expect a more optimized harness to lower this. Yet, it seems that lowering cost (shifting curves horizontally to the left) would not dramatically change the expenditure horizon or the maximum speedup achieved.
GPT-5 and Opus-4.1 did not make any meaningful progress at all. In particular, raw agent trajectories overstate progress due to statistical noise. The other four models after re-validation have positive expenditure horizons ranging from $600 to $3.3K.
Agents made some reasonable contributions that map to roughly 1-1.5% in speedup (equivalent to 1-2 human contributions). We think the other contributions are either not real (e.g., do not hold up when re-validated) or low-quality (e.g., brittle tuning that is overly curve-fit to the exact loss target). We discuss later the mergeability of their contributions as judged by the speedrun’s maintainer.
Overall, although some models have expenditure horizons in the thousands of dollars, they are small relative to the overall human labor, indicating that autonomous agent optimization has so far had minimal effect on AI R&D progress in NanoGPT.
Overall, the maintainer estimates roughly 70% of ideas would be mergeable, but what the agent mostly explores (e.g., hyperparameter tuning) is of low novelty. The estimated mergeable share of speedup, however, is smaller — roughly 60% for Opus-4.8 and 50% for GPT-5.5.
Some model-generated ideas seem genuinely good.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does AI-assisted research sacrifice exploration breadth for productivity gains?- How much faster and cheaper are AI agents compared to human researchers?
- What distinguishes artifact efficiency improvements from research process efficiency improvements?
- Do gains in optimization benchmark scores translate to gains in real research efficiency?
- How does rising researcher count relate to declining output per scientist?
- Can we distinguish agent effort from actual research output quality?
- How do AIDE2's held-out gains compare to matched-budget test-time search baselines?
- Can automated benchmarks fairly evaluate messy real-world research tasks?
- What biases affect how we measure progress on research leaderboards?
- Why do estimates for task-level performance differ so much from full job automation timelines?
- What real-world tasks most clearly expose gaps between benchmark performance and actual capability?
- Can benchmark scores on verifiable tasks transfer to unseen problems outside the training domain?
- How does automated R&D affect the efficiency of the research process itself?
- Does AI research acceleration compound into faster field-wide progress over time?
- How do technological spillovers between research sectors compound growth rates?
- How much sector-level productivity spillover does real AI research exhibit?
- What drives the gap between AI capability and actual cost savings in practice?
- Why do AI benchmarks show rapid saturation from near-zero to near-perfect?
- How do open-world evaluations correct distortions that automated benchmarks introduce?
- Why do static benchmarks miss frontier capabilities that open-world tasks reveal?
- How should evaluation frameworks account for the computational cost of frontier AI capability?