The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

Paper · arXiv 2609.11873 · Published September 10, 2026
Frontier AI Risk & RSI

Abstract Recursive self-improvement (RSI) enables AI systems to turn experience and feedback into persistent changes that improve both their capabilities and the process of future improvement. We first use the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, then introduce the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement. Next we examine RSI across scenarios (e.g., scientific discovery, embodied intelligence, software engineering), highlighting their distinct requirements and development speeds. Drawing on diverse industry practices and preliminary empirical evidence, we connect RSI research with practical systems and identify key challenges to achieving genuine RSI.

Introduction. Recent frontier-model development illustrates several forms of scaling in the improvement pipeline. Kimi K3 and Qwen3.8-Max contain 2.8 trillion and 2.4 trillion parameters, respectively, and each supports a context window of approximately one million tokens [1, 2]. The development process is also expanding. During the six months preceding GPT-5.6, OpenAI reports that the share of research compute devoted to internal coding inference grew 100-fold and internal agentic token use grew 22-fold, and average daily output tokens per active researcher exceeded twice the previous peak observed with GPT-5.5 [3]. Scale accumulates across training runs, model-assisted experiments, inference, evaluation, and human validation.

Despite the growing use of agentic tools and API-based automation, developers must still determine what to improve, construct the required resources, and establish whether each change works [4, 5]. As AI systems take on more demanding tasks, the cost of scaling this end-to-end development process becomes a bottleneck [6–9]. The following three challenges occur at different stages of the model lifecycle and motivate RSI:

• Development challenge 1: Resource-intensive foundation-model training. Foundation-model development remains resource-intensive across data preparation, architecture design, distributed optimization, and evaluation. Kimi K3 activates 16 of 896 experts and reports an approximate 2.5-fold improvement in scaling efficiency over Kimi K2, while Qwen3.8-Max activates 95 billion of its 2.4 trillion parameters [1, 2]. Sparse activation reduces per-token computation, but training at this scale still couples expert routing, parallelism, multimodal integration, long-context optimization, and systems design. OpenAI further reports that GPT-5.6 Sol designed and ran hundreds of experiments on its speculative-decoding draft model and monitored training through hardware failures and instability. The resulting changes improved token-generation efficiency by more than 15% [10]. Data quality also matters in addition to quantity. OpenAI’s GDPval illustrates that its 1,320 professional tasks required roughly 9,240 expert-hours in total, with contributors averaging more than 14 years of experience [6]. The Humanity’s Last Exam pipeline logged more than 70,000 submission attempts and sent approximately 13,000 model-stumping questions to expert review before producing a 3,000-question benchmark [7]. Architecture search remains expensive because candidate structures interact with their data, optimization, and hardware regimes. AgentNAS addresses part of this bottleneck by using an LLM to propose a task-specific seed architecture and construct its search space, but candidate selection still depends on combinatorial search under an externally specified objective [11].

• Development challenge 2: Scaling feedback and learning environments. Synthetic data and reinforcement learning automate parts of capability development, but introduce substantial requirements for generating and evaluating experience, requiring both experience-generation infrastructure and reliable mechanisms for evaluating and retaining updates. For instance, DeepSeek-V3.2 reports a post-training computational budget exceeding 10% of its pretraining cost [12], while NVIDIA’s AIMO-2 pipeline generated 3.2 million long-reasoning solutions and 1.7 million tool-integrated solutions, in addition to curating 540,000 problems [4].

• Development challenge 3: Recurring adaptation after deployment. Deployed systems consistently introduce changing documents, unfamiliar tools, incomplete context, and workflows involving interdependent actions. Improving such systems requires engineers to manually diagnose failures, revise retrieval and tool interfaces, manage persistent state, and repeat regression testing through discrete, human-led releases. Anthropic reports that agentic workloads use approximately four times as many tokens as ordinary chat, rising to about fifteen times for multi-agent systems because of longer contexts, coordination, environment setup, and end-to-end verification [13], while Meta reports that FBDetect identifies thousands of infrastructure regressions each week, and diagnosing one such regression required roughly ten engineer-hours [5].

The burdens arise because model improvement remains a sequence of costly, externally coordinated interventions. To address these barriers, RSI inspects whether part of that coordination can become a persistent capability of the system being improved. We define recursive self-improvement (RSI) as an autonomous, closed-loop process in which an AI system identifies its own limitations, develops and validates improvements, and uses the resulting capabilities to improve the improvement process itself. The model evolution paradigm of RSI spans three dimensions: autonomy, efficiency, and innovation.

Related work. Existing surveys provide complementary taxonomies of self-evolving systems and the mechanisms from which improvement loops are built. These works establish much of the technical vocabulary on which our analysis relies. Our survey differs in four respects.

• Improvement loop as the unit of analysis. Prior surveys organize work by stages of self-evolution, update objects, timing, or technical mechanisms [26–29]. These views explain what changes and how the change is produced, but systems that update the same component may assign very different decisions to AI. We trace a complete loop: what triggers improvement, who proposes and validates a change, what persists, and which later decisions use the retained change.

• Responsibility as the autonomy criterion. Related frameworks examine capability levels, co-evolution, dynamic agent state, and AI-for-AI systems [30–33]. We operationalize autonomy through the improvement decisions transferred from external designers to the AI rather than through model capability or the number of automated components. Our five levels distinguish responsibility for execution, strategy selection, experience acquisition, environmental adaptation, and recursive inheritance.

• Separate evidence for recursion and performance. Evaluation surveys study model-based judgment, agent assessment, rubric-guided learning, and oversight failures [34–38]. Higher task performance alone does not show that an improvement mechanism was revised, retained, and reused. We distinguish structural recursion, in which a revised improvement mechanism governs a later round, from effective recursion, in which that mechanism produces stronger successors under comparable budgets and independent evaluation.

• Mechanisms compared across operating conditions. Work on correction, synthetic data, lifelong learning, memory, prompt optimization, and workflow design explains how individual components improve [39–49]. We examine how these mechanisms function within complete loops across science, embodied intelligence, software engineering, healthcare, and industrial practice, where feedback cost, validation, and external control differ [50–52]. Comparing the same loop questions across these settings reveals when a method depends on cheap executable feedback, repeated interaction, expert review, or production infrastructure.

These choices together position the survey between a catalog of improvement mechanisms and a general hierarchy of AI capabilities. Our aim is to determine which parts of an improvement loop current systems can assume, how retained changes affect later rounds, and what evidence supports claims of recursive progress.

Method. Figure 1 maps representative systems across this progression, while Figure 2 isolates the corresponding loop structures. At each level, we identify where the improvement loop closes, what is retained for later rounds, and which critical decisions remain under human control, then introduce the techniques that implement this division of responsibility.

• (L1) Improvement Execution Autonomy. Humans specify what should be improved, how it should be improved, and what constitutes success, while AI executes candidate updates. For example, FineWeb-Edu uses a model to apply human-defined educational-quality labels across a web corpus without choosing the labeling criterion [22].

• (L2) Improvement Strategy Autonomy. The objective, task boundary, and evaluation criteria remain externally fixed, but AI diagnoses weaknesses and decides how to improve the system. For example, Self- Harness uses execution traces to propose and test edits to its agent harness under a fixed benchmark and promotion rule [23].

• (L3) Learning-Signal or Experience-Acquisition Autonomy. The system also determines the experience needed for its next improvement round. For example, SIMA 2 uses assessments of current behavior to generate later practice tasks that target observed skill weaknesses [24].

• (L4) Environment Adaptation Autonomy. The improvement loop uses deployment interaction to revise persistent system state under external acceptance and governance rules. For example, PANDO admits or demotes reusable rules during a long-running interaction according to observed outcomes, so later actions inherit earlier experience [25].

• (L5) Recursive Inheritance Autonomy. The system persistently revises a mechanism that governs subsequent improvement, such as an improver, verifier, or successor-generation procedure. For example, A-Evolve-Training revises its research policy after development scores fail to predict external gains and uses the revised policy to direct the next training round [17].

While the autonomy levels describe the structure of an improvement loop, their practical meaning depends on the feedback available in a domain, as the same retained update may be straightforward to test in software engineering and difficult to validate in a physical or clinical setting. We consider science, embodied intelligence, software engineering, and healthcare because they expose four distinct feedback regimes, including experimental evidence with uncertain attribution, physical interaction with costly trials, executable tests with incomplete specifications, and high-stakes outcomes under expert oversight. These regimes allow us to compare how feedback cost and reliability affect the retention and reuse of improvements.

• (S1) RSI for Science. Scientific discovery involves open-ended exploration, costly experiments, and feedback that may not clearly identify the source of failure. We examine how accumulated evidence can improve scientific hypothesis modules, experimental agents, and reflection or improvement mechanisms, with attention to whether these changes support subsequent research beyond the current scientific result.

• (S2) RSI for Embodied Intelligence. Embodied agents generate experience through their own actions, while failures may arise from interacting perception, planning, and control components. Physical trials also impose limits on exploration and repeatability. We examine the evolution of environments and curricula, skills and agent harnesses, policies and action models, and world models and evaluators, focusing on how interaction feedback supports validated improvements that can be reused in later tasks.

• (S3) RSI for Software Engineering. Software engineering makes both the developed artifact and the developing agent accessible to executable modification and testing. We examine how repository feedback supports persistent changes to coding-agent implementations and harnesses, development experience and collaboration, and the improvement process itself.

Discussion. • Observation 1: Frontier gains differ in magnitude and timing. By 2026, advanced mathematics and graduate-level science reach HCI values of 86.4 and 85.8, while broad knowledge reaches 77.2. The corresponding values for legal reasoning, multimodal reasoning, and frontier academic breadth are 64.5, 62.2, and 60.4. The annual increment ∆Td,y = Td,y −Td,y−1 exposes different temporal patterns. Broad knowledge rises by 32.8, 26.9, and 17.6 points across the three annual transitions, indicating steady but slowing headroom closure. Legal reasoning similarly slows from a 48.2-point gain in 2024 to 11.4 in 2025 and 4.9 in 2026. Advanced mathematics follows a different path, whose increment grows from 32.8 points in 2025 to 53.6 in 2026. Multimodal reasoning gains 59.7 points in 2025 but only 2.5 in 2026. A single aggregate benchmark would conceal these differences in both level and trajectory shape.

• Observation 2: Interactive capabilities retain larger gaps. Software engineering reaches an HCI of 52.6 in 2026, search and terminal agents reach 56.8, and tool agents reach 39.9. Relative to graduate-level science at 85.8, their normalized headroom closure is lower by 33.2, 29.1, and 45.9 points, respectively. Tool agents improve sharply in 2026, from 8.2 to 39.9, yet remain the lowest trajectory. Software engineering gains 40.8 points in 2025 and 11.9 in 2026. The leading cybersecurity-agent trajectory reaches 91.9, producing a 52.0-point difference from tool agents. This comparison requires caution because the later Cybench observations use changed task subsets or pass@1 aggregation, as indicated by the dashed line. Even with that qualification, the figure shows that gains in bounded or readily verified environments have not transferred uniformly to long, stateful workflows. Model developers increasingly turned their attention to agentic coding and tool use after 2024, and these capabilities became prominent research and engineering targets during 2025 and 2026 [25, 59–61]. Such tasks require planning, environment-state tracking, tool selection, result interpretation, and revision of subsequent actions. Errors propagate across the trajectory, so data collection and evaluation must cover complete interactions. In the paper’s autonomy taxonomy, bounded evaluations mainly exercise L1–L2 capabilities, environment tasks increasingly require L2–L3 capabilities, and interactive workflows expose the verification, memory, and adaptation requirements associated with L3–L4 operation.

• Observation 3: Remaining headroom concentrates the potential value of RSI. The hatched post-2026 region illustrates how persistent and verified recursive self-improvement could preferentially affect domains with larger remaining gaps. To express this relationship consistently, the illustrative endpoint for domain d is Software engineering and tool use are therefore particularly relevant to RSI. Human-led updates require repeated environment construction, trajectory collection, failure diagnosis, training or harness revision, and regression testing as interfaces and repositories change. The methods reviewed later generate practice from observed weaknesses [61], distill trajectories into reusable rules and tools [25, 62], and retain tested harness or code changes for future tasks [23, 63]. These mechanisms can direct successive updates toward failures observed during deployment and reduce the capability gaps represented by the illustrative extensions.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI research automation sustain progress through accelerating feedback loops? What limits recursive self-improvement in autonomous AI systems? Do evolved harnesses learn transferable strategies or task-specific optimization artifacts? Can code harness improvements rival direct model scaling for capability? Do individually safe AI actions create unsafe outcomes in integrated systems? How do real-world evaluations reveal AI capabilities that benchmarks hide? Should governance of agentic AI systems be runtime or design-time? Can language models reliably simulate personas and predict behavior? What causes coordination failures in multi-agent language model systems?