SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

How does control over improvement decisions scale in AI systems?

What are the distinct ways that responsibility for improvement decisions can transfer from human designers to AI systems, and how does this relate to measurable progress in different domains?

Synthesis note · 2026-10-08 · sourced from Frontier AI Risk & RSI

The paper proposes a five-level taxonomy for recursive self-improvement (RSI), built on a criterion it calls responsibility: not capability or the number of automated components, but "the improvement decisions transferred from external designers to the AI." The levels run (L1) Improvement Execution Autonomy, where "humans specify what should be improved, how it should be improved, and what constitutes success, while AI executes candidate updates" — its example is FineWeb-Edu applying human-defined quality labels; (L2) Improvement Strategy Autonomy, where objective and evaluation stay fixed but "AI diagnoses weaknesses and decides how to improve the system," exemplified by Self-Harness proposing and testing its own harness edits "under a fixed benchmark and promotion rule"; (L3) Experience-Acquisition Autonomy, where the system "determines the experience needed for its next improvement round," exemplified by SIMA 2 generating practice tasks from observed skill weaknesses; (L4) Environment Adaptation Autonomy, where deployment interaction revises "persistent system state under external acceptance and governance rules," exemplified by PANDO admitting or demoting rules by observed outcome; and (L5) Recursive Inheritance Autonomy, where the system "persistently revises a mechanism that governs subsequent improvement, such as an improver, verifier, or successor-generation procedure," exemplified by A-Evolve-Training revising its own research policy.

The paper ties this ladder explicitly to verifiability. Its Observation 2 reports that "interactive capabilities retain larger gaps" on its Headroom-Closed Index: software engineering reaches an HCI of 52.6 and tool agents 39.9 in 2026, far below graduate-level science at 85.8, even though tool agents jumped from 8.2 to 39.9 within the year. The paper reads this gap through its own levels: "bounded evaluations mainly exercise L1–L2 capabilities, environment tasks increasingly require L2–L3 capabilities, and interactive workflows expose the verification, memory, and adaptation requirements associated with L3–L4 operation." The levels that require the system to supply or validate its own feedback are the levels where measured progress lags, because "gains in bounded or readily verified environments have not transferred uniformly to long, stateful workflows."

Measured against What separates self-improvement from policy improvement?, this is the same responsibility-transfer criterion cast as an ordinal ladder rather than two binary dials: that paper asks only whether the improver sits inside the agent and whether the standard is external, while this one splits "improver inside the agent" into four successively deeper commitments. It also gives internal structure to the line drawn in Are self-refinement and recursive self-improvement actually the same thing?: L1–L2, fixed objective and externally validated, sit on that survey's bounded side, while L3–L5, self-supplied experience and self-revised governing mechanisms, sit on its open-ended side. The paper's own L2 example, Self-Harness, is the mechanism Does harness self-improvement memorize tasks instead of learning broadly? examines for overfitting risk, and its account of humans retaining "what should be improved" while AI executes echoes the split How should AI agents and humans divide research tasks? measures empirically in one lab's own R&D records.

The excerpt states the taxonomy and names one example system per level; it does not report how reliably real systems can be assigned to a level, nor show any system reaching L4 or L5 outside its single cited case each. The HCI figures behind Observation 2 carry the paper's own caveat that "the later Cybench observations use changed task subsets or pass@1 aggregation," a comparability caveat that bears on the broader claim that verifiability predicts RSI progress. What the excerpt supports is a vocabulary for locating a system's autonomy and a self-reported correlation between verification difficulty and headroom closure, not proof that climbing the five levels is a mechanism rather than a post-hoc description, or that the ordering is causal.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How should humans and AI agents share control and decision-making? How can humans maintain effective oversight as AI systems scale? How does AI adoption reshape collaboration patterns in knowledge work?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 92 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

five autonomy levels mark recursive self-improvement by who controls each improvement decision — from execution to recursive inheritance