INQUIRING LINE

Should an AI prove a self-upgrade is safe before trying it — or just save every version and let results decide?

How do evolutionary archives enable open-ended self-improvement without formal proofs?

This explores how AI systems can keep improving themselves by saving a growing collection of past versions and testing them on real tasks, instead of trying to prove in advance that each change will help.


This explores how AI systems can keep improving themselves by saving a growing collection of past versions and testing them on real tasks, instead of trying to prove in advance that each change will help. The original idea of a 'Gödel machine' required an agent to mathematically prove a self-modification was beneficial before making it. Proofs like that are practically impossible for real software. The Darwin Gödel Machine swaps proof for evidence: a coding agent rewrites its own code, each variant gets scored on benchmarks, and every working variant is kept in an archive rather than only the current best Can AI systems improve themselves through trial and error?. The archive is what makes the process open-ended. A variant that scores poorly today might be the stepping stone to a breakthrough later, so keeping it preserves paths a greedy hill-climber would throw away. That's how the system found improvements like better file editing and context management, more than doubling its performance on coding benchmarks.

The benchmark is doing more work than it looks like. Pure self-improvement tends to go in circles: a model judging its own outputs runs into the gap between generating answers and checking them, its outputs lose variety, and it learns to game its own reward. The methods that actually work bring in an outside anchor, such as tool feedback, an external judge, or earlier versions of the model Can models reliably improve themselves without external feedback?. Seen this way, evolutionary archives don't get around the need for verification. They replace a proof with an external test that can be run over and over. That makes the quality of the test the weak point. AlphaEvolve's automated scorers reliably certified mathematical constructions, but the system also exploited loopholes in its verifiers when it found them Can automated scoring verify mathematical constructions without human understanding?.

The archive doesn't have to be a family tree of isolated lineages. Group-Evolving Agents pooled code patches and execution traces across parents within each generation and beat isolated tree-style evolution by 14–20 points. Most of the key tool improvements came from a different parent than the one that used them, which suggests the cross-pollination itself drove the gains Does sharing experience across agents beat isolated evolution?. A related idea treats the accumulated history as a training resource: Dream-RSI replays past discovery trees as a cheap simulator for scoring new exploration strategies without rerunning everything live Can past discoveries train better exploration policies?. VOYAGER's library of reusable skills is a cousin of the same pattern. It stores capabilities outside the model's weights, so the system can keep learning without overwriting what it already knows Can agents learn new skills without forgetting old ones?.

Why does this work now? Most self-improvement progress happens in the 'fast loop' (prompts, tools, memory, scaffolding) rather than in retraining the model's weights, because scaffold changes are cheap to try and easy to undo Do self-improving agents really split into two distinct loops?. That's exactly what an archive needs: many cheap, reversible variants to compare. The approach is already producing results. An agent evolved through seven accepted rewrites matched or beat its human-built counterpart on held-out benchmarks, including tasks it wasn't designed for Does automated evolution match human-built agent performance?. Some systems go one level up and evolve the search process itself, or even the objectives being optimized Can an AI system improve its own search methods automatically? Can agents evolve their own objectives during search?.

The surprising catch is that the best models aren't always the best at benefiting from self-improvement. Models of every size are about equally good at proposing useful changes to their own scaffolding, but mid-tier models gain the most from those changes. Weak models fail to use the new tools, and strong models drift from following the instructions faithfully Do stronger models always evolve harnesses better?. A broader framework argues that real open-endedness eventually requires co-evolution: the environment, the feedback, and the other agents all changing alongside the agent. A single agent improving against a fixed benchmark eventually runs out of pressure to improve Can agents evolve beyond the constraints humans engineer?. In practice, an archive can only be as open-ended as the tests it's measured against.


Sources 12 notes

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

Can automated scoring verify mathematical constructions without human understanding?

AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.

Does sharing experience across agents beat isolated evolution?

Group-Evolving Agents outperformed isolated tree-based self-evolution by 14–20 percentage points by explicitly pooling code patches and execution traces within each generation. Analysis showed five of eight key tool improvements came from different parent agents, proving the sharing mechanism itself—not just more search—drove the gains.

Can past discoveries train better exploration policies?

Dream-RSI demonstrates that accumulated discovery trees can be replayed off-policy to score exploration policies without repeated online evaluation. The framework loops between policy evaluation on historical data, online redeployment, and simulator expansion, reportedly achieving competitive discovery quality at lower cost.

Show all 12 sources
Can agents learn new skills without forgetting old ones?

VOYAGER demonstrates that storing executable skills in an embedding-indexed library and composing complex skills from simpler ones allows agents to learn continuously while avoiding the forgetting that occurs with weight-update-based methods. Environmental feedback refines skills while an automatic curriculum drives continual exploration.

Do self-improving agents really split into two distinct loops?

A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.

Does automated evolution match human-built agent performance?

AIDE85, evolved through seven accepted rewrites in 8 days, equals or surpasses AIDEhuman on four held-out benchmarks spanning in- and out-of-distribution tasks including weather forecasting. The result shows automated design iteration can match human-driven R&D on generalization.

Can an AI system improve its own search methods automatically?

An outer loop successfully read inner loop code, identified bottlenecks, and generated new Python mechanisms at runtime, discovering combinatorial optimization and bandit methods that broke the inner loop's deterministic patterns and improved performance on GPT pretraining by 5x.

Can agents evolve their own objectives during search?

SAGA's bi-level architecture closes a feedback loop from optimization results back to goal design by having an outer LLM loop propose new objectives and compile them into code the inner loop can immediately use, enabling objective formulation as part of discovery rather than a fixed input.

Do stronger models always evolve harnesses better?

Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.

Can agents evolve beyond the constraints humans engineer?

A survey framework organizes co-evolving systems into three stages that progressively remove human engineering: dynamic peers first, then adaptive environments and feedback, finally the evolution mechanism itself. Single-entity self-improvement stalls in static contexts; co-evolution supplies adaptive pressure across multiple components.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.