INQUIRING LINE

Counting AI's computing power looks like a checkable way to pace frontier AI, but does more compute still reliably mean more capability?

What makes compute allocation a verifiable lever for pacing frontier AI development?

This explores why the amount of computing power used to train and run AI systems might work as a measurable, checkable way to slow or pace frontier AI development. The corpus doesn't directly cover compute governance, but it has a lot to say about whether compute is still a reliable stand-in for capability.


This explores why compute, meaning the chips and processing power that go into building and running AI, is often proposed as a measurable way to pace frontier AI. The basic idea is that compute is physical, countable and concentrated, so it seems easier to verify than something abstract like 'capability.' Be aware that this collection has no notes on compute governance itself: nothing on chip tracking, training-run thresholds or reporting regimes. What it does have is research that tests the assumption underneath the whole idea, which is that more compute reliably means more capability. That evidence is more complicated than the policy framing suggests.

The first problem is that capability is moving away from raw training scale. A 3B-parameter model reached frontier-level math and coding scores mainly through careful post-training design, not size Can small models match frontier reasoning without massive scale?. Related work finds that how a model is trained matters more than how much compute it uses at inference time Can non-reasoning models catch up with more compute?. If recipes matter as much as raw processing power, counting floating-point operations misses much of what drives progress.

The second problem is that compute is split between training and use, and the two can stand in for each other. Snell et al. showed that a smaller model given more 'thinking' compute at answer time can match a larger model on hard prompts Can inference compute replace scaling up model size?. Spreading that same compute adaptively, giving more to harder prompts, does better still Can we allocate inference compute based on prompt difficulty?. A rule that only counts training runs could miss capability gained after deployment. The same pattern now applies to reward models, which improve by reasoning before they score Can reward models benefit from reasoning before scoring?. Taken further, breaking a task into very small steps and voting on each one let small non-reasoning models complete million-step tasks without errors Can extreme task decomposition enable reliable execution at million-step scale?. How a system is organized can replace how big it is.

The third problem matters most for pacing: systems that improve themselves. The Darwin Gödel Machine improved its own coding-agent design by trial and error, more than doubling its SWE-bench score Can AI systems improve themselves through trial and error?. On long optimization tasks, success depends less on starting quality than on persistence, meaning how long an agent keeps running the benchmark-edit-retry loop What predicts success in ultra-long-horizon agent tasks?. In that setting, compute turns into wall-clock time spent iterating, which is a different thing to measure than a training run's size. The good news is that current research agents mostly recombine known techniques rather than invent new ones Do frontier AI agents actually conduct novel research or just optimize?.

What the reader might not expect is that verifying compute is the easy half. A compute limit only paces development if you can also check what capability a given amount of compute buys, and standard benchmarks both overstate and understate that. One proposed fix is open-world evaluations of messy, long tasks with costs reported explicitly Do automated benchmarks hide what frontier AI systems can really do?. That pairing is the real implication here: compute can be checked, but it only works as a pacing tool if capability evaluation keeps up. For the policy question itself, the collection would need sources on compute governance that it doesn't have yet.


Sources 10 notes

Can small models match frontier reasoning without massive scale?

A 3B model trained with curriculum SFT and multi-domain RL reaches 94.3 AIME26 and 80.2 LiveCodeBench scores matching much larger systems. The result is bounded to verifiable tasks with checkable ground truth, where RL can provide clean reward signals.

Can non-reasoning models catch up with more compute?

Reasoning models persistently outperform non-reasoning models regardless of inference budget because training instills a reasoning protocol that makes additional tokens productive. The gap is fundamentally about deployment mechanisms and training structure, not raw capability.

Can inference compute replace scaling up model size?

Snell et al. (2024) showed that inference-time compute trades off against model parameter scaling, especially on difficult prompts. This reveals pretraining and inference compute are not independent resources.

Can we allocate inference compute based on prompt difficulty?

Research shows inference effectiveness varies dramatically by prompt difficulty. Reallocating the same total compute adaptively—giving easy prompts less and hard ones more—substantially outperforms larger models under uniform budgets.

Can reward models benefit from reasoning before scoring?

Three independent teams (RRM, RM-R1, DeepSeek-GRM) discovered that adding chain-of-thought reasoning before reward scoring enables adaptive test-time compute scaling for evaluation. Reasoning-based approaches raise the capability ceiling of reward models beyond what outcome-based evaluation achieves.

Show all 10 sources
Can extreme task decomposition enable reliable execution at million-step scale?

MAKER solves million-step tasks with zero errors by decomposing into minimal subtasks, applying voting at each step, and flagging correlated errors. Surprisingly, small non-reasoning models suffice when decomposition is extreme enough, inverting the standard approach to hard problems.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

What predicts success in ultra-long-horizon agent tasks?

Across 17 frontier models on 36 expert-curated optimization tasks, repeated benchmark-edit-incorporate cycles within a wall-clock budget proved the dominant success predictor. Most models terminated early or burned budget unproductively; Claude Opus 4.6 stood out as persistent.

Do frontier AI agents actually conduct novel research or just optimize?

Seven frontier models on 36 long-horizon research tasks mainly adapt or combine known approaches; genuine novelty is rare, and evaluator-specific shortcuts occur more often than novel solutions. Performance varies substantially across runs.

Do automated benchmarks hide what frontier AI systems can really do?

Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.