What actually drove the nanogpt speedrun's massive gains?
A breakdown of the nanogpt speedrun's 31x speedup asks whether most progress came from deep invention or from adapting and importing existing ideas. The answer matters for understanding what AI R&D acceleration really means.
METR researcher Manish Shetty measured AI R&D progress using the nanogpt speedrun, a public leaderboard where 36 contributors submitted 77 records between May 2024 and March 2026, cutting training time on a small GPT-2 pretraining task from 45 minutes to 1.43 minutes, a "31x speedup." Breaking that cumulative gain down by the nature of each contribution, Shetty finds imported ideas drove 6.7x, adapted ideas 3.0x, and invented ideas only 1.6x — so, as he puts it, "relatively shallow contributions have had a big impact," with "shallow and moderate contributions together" driving "roughly 21x of the 31x." Record 12, which simply migrated to the newly released FlexAttention library, got a 30% speedup on its own.
The decomposition is read against time. Early records (the first 20, May 2024–Jan 2025) were "mostly imported or adapted to catch up to frontier research" — RoPE, ReLU2, QK-norm, FlexAttention — with only one invention, the Muon optimizer (record 3), which later reached frontier-scale models GLM-4.5 and Kimi K2. Later records (Jan 2025–Mar 2026) "increasingly invented new ideas," rising to 33% invented, including record 58's Paired Head Attention and record 62's Bigram Hash Embedding. So invention became proportionally more common even as it contributed the smallest share of total speedup. The same pattern holds for the four AI-agent-credited records (Hiverge, Locus, Aster, Station, late 2025–early 2026): all are "real improvements," none reached "the deep or breakthrough end of the scale," and Station's catch of a silently-ignored config bug is itself a shallow-but-real fix humans had missed for months.
This sits in tension with Is AI development already being handed to AI systems?: where Anthropic's own speedup figures (3x to 52x) support an RSI-is-coming-soon argument, Shetty's independent decomposition of a comparable speedup finds its substance is mostly import and adaptation, not invention — a more skeptical read of what "speedup" evidence actually shows. It extends Do frontier AI agents actually conduct novel research or just optimize? by showing the same engineering-optimizer pattern — composing established techniques rather than inventing — describes most human contributions too, not just agents; the gap Shetty observes between agents and humans may be a gap in disclosure and task framing more than in kind. It parallels Does Sakana's AI Scientist deliver autonomous research without human help? as another independently run check that complicates self-reported claims of AI research acceleration. And it bears on Do fixed-budget efficiency gains translate to real research progress?: Shetty names the same interpretive gap, listing contamination, survivorship bias, scale-dependence, and composability as open problems in reading leaderboard progress as R&D-trend evidence.
The excerpt is explicit about its limits: GPT-2-scale pretraining on 8×H100s "is not frontier pretraining," the leaderboard shows only ideas that worked (survivorship bias), and how pretraining acceleration composes with the rest of the AI R&D stack is "an open question." The four agent records carry no disclosure of inference compute spent, failed attempts, or how much a human co-contributor directed the run, so, per Shetty, these contributions "are hard to interpret on their own." The finding therefore supports a narrower claim than "AI agents are approaching human-level AI research": it supports only that, on this one small-scale, well-documented leaderboard, both humans and the few recorded agent contributions have so far produced mostly shallow-to-moderate gains, with deep invention present but a minority contributor to total speedup.
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do frontier AI agents actually conduct novel research or just optimize?
Exploring whether current long-horizon research agents generate genuine methodological novelty or primarily recombine established techniques. This matters for understanding how close we are to recursive self-improvement through AI.
extends this pattern to humans: most of the speedrun's cumulative gain is composition, not invention, too
-
Is AI development already being handed to AI systems?
Anthropic reports rising task length, code authorship, and speedup metrics as evidence that AI systems are taking on development work. The question is whether these measures actually demonstrate autonomous delegation of R&D or reflect improvements in assisted productivity.
contrasts Anthropic's speedup framing with a decomposition showing most gain is import and adaptation
-
Does Sakana's AI Scientist deliver autonomous research without human help?
Can an AI system truly run the complete research lifecycle alone, or does it still need human guidance and oversight? This matters for understanding whether automated research can scale.
parallel independent check complicating self-reported claims of AI research acceleration
-
Do fixed-budget efficiency gains translate to real research progress?
The paper measures research efficiency as optimization gains under a fixed evaluation budget, but this differs from the real-world costs of R&D spending and human effort. Does this narrower measurement actually predict whether AI agents reduce the true cost of research discovery?
names the same gap between leaderboard progress and research-efficiency evidence
-
How much progress have AI agents actually made on NanoGPT?
METR compares AI agent optimization to human researcher productivity on a popular benchmark. By measuring where their improvement curves intersect, they ask whether autonomous systems are meaningfully accelerating AI R&D or mostly chasing noise.
Evidence for: a different METR metric finds agentic optimization added minimal effect, consistent with imported ideas dominating the speedup
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Research note: Evidence on AI R&D Progress from NanoGPT
- Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT
- How much does AI impact development speed? An enterprise-based randomized controlled trial
- Summary of METR's predeployment evaluation of Claude Opus 5.5
- We are Changing our Developer Productivity Experiment Design
- Reasoning Structure of Large Language Models
- Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
- Progress Measures For Grokking Via Mechanistic Interpretability
Original note title
METR finds shallow contributions drove most of the nanogpt speedrun's 31x speedup despite rising invention over time