What really makes AI 'self-improvement' different from ordinary training — who's doing the improving, or who decides what counts as better?
What separates an internal improver from an external improvement standard?
This explores the difference between who does the improving (a system changing itself versus being changed from outside) and who decides what counts as better (a standard the system sets for itself versus one that comes from outside it).
This explores two separate questions that usually get blurred together when people talk about AI improving itself: who does the improving, and who decides what 'better' means. The clearest framing in the corpus treats these as two independent dials What separates self-improvement from policy improvement?. One dial asks whether the improver sits inside the agent or outside it. The other asks whether the performance standard comes from the agent itself or from somewhere else. Under this view, 'recursive self-improvement' and ordinary iterative training are not different kinds of thing. They are the same improvement cycle with the dials set differently. The surprising lesson is that the second dial matters far more than the first.
The reason is circularity. If a model judges its own progress by its own standard, it can only find errors it is already able to see. That limit has a name: the generation-verification gap. A model can improve itself only to the extent that it checks answers better than it produces them, and on factual tasks that gap disappears entirely What limits how much models can improve themselves?. A broader review argues that pure self-improvement stalls for this reason, along with diversity collapse and reward hacking. It also finds that the methods that do work quietly bring in an outside anchor: an earlier model version, a third-party judge, user corrections, or feedback from tools Can models reliably improve themselves without external feedback?. So an internal improver is fine. An internal standard is where things go wrong.
Revising reasoning shows this most clearly. When reasoning models revise their own uncertain answers, they mostly keep wrong answers, and smaller models often switch correct answers to incorrect ones Does self-revision actually improve reasoning in language models?. Give the same revision step a critique from an external model and accuracy goes up. The act of revising is identical in both cases. What changes the outcome is where the judgment comes from Does revising your own reasoning actually help or hurt?. The same split appears in test-time scaling. Internal methods train a model to reason more on its own, while external methods add search and verification at inference time. The two work together rather than compete How do internal and external test-time scaling compare?.
The most interesting work sits on the line between the two dials: what if the standard improves too? Meta-Rewarding adds a 'meta-judge' that grades the judge, so the actor and its evaluator get better together. It pushed past the plateau a fixed judge hits, without human supervision Why do self-improvement loops plateau without updating the judge?. The Red Queen Gödel Machine makes evaluation part of the improvement loop itself. That lets self-improvement reach tasks like writing and proofs, which have no fixed grader Can evaluators improve alongside the agents they score?. These systems blur the inside/outside distinction. Whether they avoid circularity or just move it up one level is the open question. ModularRSI takes the opposite approach. It deliberately keeps the standard external by improving harness modules on data separate from the benchmark, so gains are less likely to be fitted to the test Can harness modules improve separately from benchmark data?. Sakana reports that the Darwin Gödel Machine's self-discovered improvements carried over to other models and other programming languages. That transfer is a kind of outside check that the improvements were real Do agent improvements discovered in one model transfer to others?.
This distinction also shapes how risk is assessed. One survey argues that today's industrial 'self-refinement' is bounded and checked against outside measures. Open-ended recursive self-improvement, which would need to grow its own standards, still faces grounding problems that can be measured today Are self-refinement and recursive self-improvement actually the same thing?. Even frontier assessments struggle here. OpenAI rated GPT-5.6 below its 'High' self-improvement threshold based on a single debugging metric with no reported numbers Does GPT-5.6 show meaningful self-improvement capability?. The takeaway: when you hear 'AI improving itself,' ask who holds the measuring stick. A system can safely do its own improving. Trouble starts when it also writes its own measure of what counts as better.
Sources 12 notes
Generalized Agent Iteration shows that recursive self-improvement and iterative policy improvement are instances of the same cycle, separated by whether the improver sits inside the agent and whether the performance standard comes from outside. This framework reveals that reliable improvement requires external anchoring rather than pure self-reference.
Models can only improve themselves when they verify solutions better than they generate them. This gap scales with model size but vanishes entirely for factual tasks, predicting which domains benefit from self-improvement.
Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.
Evidence from QwQ, R1, and LIMO shows most revisions retain wrong answers rather than correcting them. Smaller models frequently switch correct answers to incorrect during revision, and longer chains with more revisions correlate with lower accuracy.
Revision guided by external models improves accuracy, but a model revising its own uncertain output typically amplifies confidence in wrong answers rather than correcting them. The revision source, not the revision act itself, determines the outcome.
Show all 12 sources
Research shows test-time scaling methods split into internal (training models for autonomous reasoning) and external (inference-time search and verification). They complement rather than compete; internal builds capability while external extracts performance from existing capability.
Meta-Rewarding adds a meta-judge layer that evaluates the judge's own judgments, creating preference data for both actor and evaluator. This co-evolution improved AlpacaEval 2 from 23% to 39% and Arena-Hard from 21% to 29% without supervision.
Red Queen Gödel Machine makes evaluation part of the improvement loop, allowing agents to optimize writing and proof generation without a static verifier. Co-evolved systems match fixed-evaluator performance while using fewer tokens, suggesting shared learning drives efficiency.
ModularRSI evolves harness modules independently using contrastive trajectories on benchmark-disjoint data, showing consistent gains across unseen tasks and domains. The approach isolates mechanism-level improvements from task-specific adaptation by aggregating evidence across tasks before updating components.
Sakana AI's Darwin Gödel Machine discovered improvements to agent tools and workflows that transferred to different foundation models (Claude 3.5 Sonnet, o3-mini, Claude 3.7 Sonnet) and to programming languages outside its training domain (Python-trained agents improved on Rust, C++, Go), suggesting the improvements target portable agent design rather than model-specific exploits.
A 1,250-paper survey shows that bounded, evaluable self-refinement (current industrial practice) differs fundamentally from open-ended recursive self-improvement, which remains constrained by grounding requirements, collapse dynamics, and compute limits measurable today.
OpenAI's Preparedness Framework rates GPT-5.6 Sol and Terra as High capability in cybersecurity and biorisks, but below High in AI self-improvement despite measurable gains on internal research-debugging tasks. The self-improvement rating relies on a single unquantified debugging metric.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
- Self-Improvements in Modern Agentic Systems: A Survey
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- Hyperagents
- Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
- MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?