How does control over improvement decisions scale in AI systems?
What are the distinct ways that responsibility for improvement decisions can transfer from human designers to AI systems, and how does this relate to measurable progress in different domains?
The paper proposes a five-level taxonomy for recursive self-improvement (RSI), built on a criterion it calls responsibility: not capability or the number of automated components, but "the improvement decisions transferred from external designers to the AI." The levels run (L1) Improvement Execution Autonomy, where "humans specify what should be improved, how it should be improved, and what constitutes success, while AI executes candidate updates" — its example is FineWeb-Edu applying human-defined quality labels; (L2) Improvement Strategy Autonomy, where objective and evaluation stay fixed but "AI diagnoses weaknesses and decides how to improve the system," exemplified by Self-Harness proposing and testing its own harness edits "under a fixed benchmark and promotion rule"; (L3) Experience-Acquisition Autonomy, where the system "determines the experience needed for its next improvement round," exemplified by SIMA 2 generating practice tasks from observed skill weaknesses; (L4) Environment Adaptation Autonomy, where deployment interaction revises "persistent system state under external acceptance and governance rules," exemplified by PANDO admitting or demoting rules by observed outcome; and (L5) Recursive Inheritance Autonomy, where the system "persistently revises a mechanism that governs subsequent improvement, such as an improver, verifier, or successor-generation procedure," exemplified by A-Evolve-Training revising its own research policy.
The paper ties this ladder explicitly to verifiability. Its Observation 2 reports that "interactive capabilities retain larger gaps" on its Headroom-Closed Index: software engineering reaches an HCI of 52.6 and tool agents 39.9 in 2026, far below graduate-level science at 85.8, even though tool agents jumped from 8.2 to 39.9 within the year. The paper reads this gap through its own levels: "bounded evaluations mainly exercise L1–L2 capabilities, environment tasks increasingly require L2–L3 capabilities, and interactive workflows expose the verification, memory, and adaptation requirements associated with L3–L4 operation." The levels that require the system to supply or validate its own feedback are the levels where measured progress lags, because "gains in bounded or readily verified environments have not transferred uniformly to long, stateful workflows."
Measured against What separates self-improvement from policy improvement?, this is the same responsibility-transfer criterion cast as an ordinal ladder rather than two binary dials: that paper asks only whether the improver sits inside the agent and whether the standard is external, while this one splits "improver inside the agent" into four successively deeper commitments. It also gives internal structure to the line drawn in Are self-refinement and recursive self-improvement actually the same thing?: L1–L2, fixed objective and externally validated, sit on that survey's bounded side, while L3–L5, self-supplied experience and self-revised governing mechanisms, sit on its open-ended side. The paper's own L2 example, Self-Harness, is the mechanism Does harness self-improvement memorize tasks instead of learning broadly? examines for overfitting risk, and its account of humans retaining "what should be improved" while AI executes echoes the split How should AI agents and humans divide research tasks? measures empirically in one lab's own R&D records.
The excerpt states the taxonomy and names one example system per level; it does not report how reliably real systems can be assigned to a level, nor show any system reaching L4 or L5 outside its single cited case each. The HCI figures behind Observation 2 carry the paper's own caveat that "the later Cybench observations use changed task subsets or pass@1 aggregation," a comparability caveat that bears on the broader claim that verifiability predicts RSI progress. What the excerpt supports is a vocabulary for locating a system's autonomy and a self-reported correlation between verification difficulty and headroom closure, not proof that climbing the five levels is a mechanism rather than a post-hoc description, or that the ordering is causal.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How should humans and AI agents share control and decision-making? How can humans maintain effective oversight as AI systems scale? How does AI adoption reshape collaboration patterns in knowledge work?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
What separates self-improvement from policy improvement?
Does recursive self-improvement work by the same evaluate-and-improve cycle as classical policy iteration, or are they fundamentally different processes? Understanding this distinction matters for predicting which self-improving systems remain controllable.
same responsibility criterion, recast as an ordinal five-level ladder instead of two binary dials
-
Are self-refinement and recursive self-improvement actually the same thing?
The survey explores whether current AI systems using "self-X" vocabulary describe one unified phenomenon or fundamentally different processes with distinct evidence, theory, and risk profiles.
gives that survey's bounded/open-ended cut internal structure, with L1–L2 bounded and L3–L5 open-ended
-
Does harness self-improvement memorize tasks instead of learning broadly?
When agents automatically edit their own prompts and tools based on task feedback, do those improvements generalize to new domains or just fit the training tasks? This matters because overfitting at the harness level could hide real capability gains.
examines overfitting risk in Self-Harness, this paper's own cited example of L2 strategy autonomy
-
How should AI agents and humans divide research tasks?
In building its own foundation model, Atria Dawn studied how to split work between agents and human researchers. Understanding this division matters for designing effective human-AI collaboration in technical R&D.
measures empirically the human/AI split this paper's L1–L2 levels describe structurally
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
- Atria Dawn: The Dawn of Agentic Superintelligence
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
- Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
- Recursive self-improvement of AI research agents
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds
- Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
Original note title
five autonomy levels mark recursive self-improvement by who controls each improvement decision — from execution to recursive inheritance