INQUIRING LINE

Why do AI agents that rewrite their own code get better by keeping old versions around instead of just building on the latest one?

How do evolutionary archives improve on single self-modification trajectories?

This explores why self-improving AI agents seem to do better when they keep a whole population of past versions to draw from, rather than following one chain where each version simply rewrites itself.


This explores why self-improving agents do better when they keep many past versions and their experiments, rather than following one chain where each version rewrites the last. The short answer from the corpus: a single chain of self-edits tends to run in circles. An archive breaks that loop by bringing in variety and outside evidence. Pure self-improvement hits a wall for structural reasons. The model can't reliably judge its own output, its ideas grow less diverse over time, and it learns to game its own reward. The methods that do work quietly bring in some outside anchor, such as earlier model versions, a separate judge, or feedback from tools (Can models reliably improve themselves without external feedback?). An archive builds one of those anchors into the design: earlier versions and their records stay available as reference points instead of being overwritten.

The clearest evidence is about sharing, not just storage. Group-Evolving Agents beat tree-style evolution, where each lineage stays isolated, by 14–20 percentage points. They did it by pooling code patches and execution traces across all agents in each generation. Five of the eight key tool improvements came from a different parent than the one that used them. So the gain came from ideas crossing between branches, not just from searching more (Does sharing experience across agents beat isolated evolution?). A survey of co-evolving systems points the same way: one agent improving itself in a fixed setting stalls, and progress comes from several parts pushing on each other (Can agents evolve beyond the constraints humans engineer?).

The less obvious lesson is that the value of an archive includes its failures and its history. SkillOpt found that skill learning was more stable when rejected edits were kept as warnings rather than thrown away (Does constraining edits make skill learning more stable?). Reflexion shows the same idea at small scale: agents improve across attempts by storing written notes on what went wrong (Can agents learn from failure without updating their weights?). Dream-RSI goes further and treats the full record of past discoveries as a cheap simulator. New exploration strategies can be tested against that record before being run for real (Can past discoveries train better exploration policies?). Once you keep an archive, it can also become training data and a test bed.

Archives also make gatekeeping possible. With a record of past versions, a rewrite only becomes the new baseline if it passes a check. AIDE reached parity with its human-built counterpart through just seven accepted rewrites (Does automated evolution match human-built agent performance?). Ouroboros sends its self-edits through reviewed commits and reports state-of-the-art scores. However, it has no ablation, so we can't tell how much of that comes from evolution itself (Does self-editing through reviewed commits improve agent performance?). A next step is to evolve the search method itself. One bilevel system had an outer loop rewrite the inner loop's search code and found new mechanisms that improved results 5x (Can an AI system improve its own search methods automatically?).

Two cautions. First, archives make bounded, checkable self-refinement work better. They don't turn it into open-ended recursive self-improvement, which is still limited by the need for outside grounding, by collapse, and by compute (Are self-refinement and recursive self-improvement actually the same thing?). Second, the corpus has only one direct head-to-head comparison of an archive against a single lineage, the Group-Evolving Agents result. The rest is supporting evidence from nearby work, so the overall case is suggestive rather than settled.


Sources 10 notes

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

Does sharing experience across agents beat isolated evolution?

Group-Evolving Agents outperformed isolated tree-based self-evolution by 14–20 percentage points by explicitly pooling code patches and execution traces within each generation. Analysis showed five of eight key tool improvements came from different parent agents, proving the sharing mechanism itself—not just more search—drove the gains.

Can agents evolve beyond the constraints humans engineer?

A survey framework organizes co-evolving systems into three stages that progressively remove human engineering: dynamic peers first, then adaptive environments and feedback, finally the evolution mechanism itself. Single-entity self-improvement stalls in static contexts; co-evolution supplies adaptive pressure across multiple components.

Does constraining edits make skill learning more stable?

SkillOpt's ablations show that adding a textual learning-rate budget, held-out validation gate, and rejected-edit buffer (retaining failed edits as negative feedback) produces more stable and generalizable skill improvement than allowing agents to freely rewrite their own instructions.

Can agents learn from failure without updating their weights?

Reflexion demonstrates that unambiguous environmental feedback (success/failure) enables agents to write useful self-diagnoses and improve across episodes without parameter updates. The binary signal prevents rationalization, and keeping reflections uncompressed preserves their usability.

Show all 10 sources
Can past discoveries train better exploration policies?

Dream-RSI demonstrates that accumulated discovery trees can be replayed off-policy to score exploration policies without repeated online evaluation. The framework loops between policy evaluation on historical data, online redeployment, and simulator expansion, reportedly achieving competitive discovery quality at lower cost.

Does automated evolution match human-built agent performance?

AIDE85, evolved through seven accepted rewrites in 8 days, equals or surpasses AIDEhuman on four held-out benchmarks spanning in- and out-of-distribution tasks including weather forecasting. The result shows automated design iteration can match human-driven R&D on generalization.

Does self-editing through reviewed commits improve agent performance?

Ouroboros, a harness that rewrites its tools, prompts, and core through reviewed commits, achieved 86.97% on Terminal-Bench 2.1, 90.69% on OSWorld-Verified, and 0.2301 on CL-Bench. However, the paper lacks ablation studies to isolate evolution's actual contribution to these scores.

Can an AI system improve its own search methods automatically?

An outer loop successfully read inner loop code, identified bottlenecks, and generated new Python mechanisms at runtime, discovering combinatorial optimization and bandit methods that broke the inner loop's deterministic patterns and improved performance on GPT pretraining by 5x.

Are self-refinement and recursive self-improvement actually the same thing?

A 1,250-paper survey shows that bounded, evaluable self-refinement (current industrial practice) differs fundamentally from open-ended recursive self-improvement, which remains constrained by grounding requirements, collapse dynamics, and compute limits measurable today.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.