If AI agents evolve side by side, pushing and testing each other, do they eventually design themselves — no human engineer needed?
Does co-evolution between peer agents reduce reliance on human design?
This explores whether letting AI agents evolve alongside each other, as well as alongside their environments and goals, can take over design work that human engineers would otherwise do by hand.
This explores whether letting AI agents evolve alongside each other, as well as alongside their environments and goals, can take over design work that human engineers would otherwise do by hand. The corpus mostly says yes, but in stages. One survey framework lays out three steps. First, agents get dynamic peers. Next, the environment and the feedback adapt too. Last, the evolution mechanism itself starts to evolve Can agents evolve beyond the constraints humans engineer?. The core argument is that a single agent improving itself in a fixed setting eventually stalls, because nothing new pushes back on it. Peers supply that push. Peer co-evolution is the first rung of that ladder, not the whole climb.
The most concrete evidence that peers matter comes from Group-Evolving Agents. Agents that pooled code patches and execution traces within each generation beat isolated self-evolving lineages by 14–20 points. Five of the eight key tool improvements came from a different parent than the one that used them, so the gain came from the sharing itself, not just from running more search Does sharing experience across agents beat isolated evolution?. SkillClaw does something similar at the level of users: it gathers interaction histories across many people and lets an automated evolver refine shared skills that a human would otherwise have to curate How can agent systems share learned skills across users?. In game settings, agents trained against a varied mix of co-players learn to cooperate without anyone hardcoding cooperative rules. Being exposed to exploitation by each other pushes them toward cooperation Can agents learn cooperation by adapting to diverse partners?.
A quieter lesson is that "peers" don't have to be other agents. POET pairs agents with environments that evolve alongside them and lets solutions move between problems. It solved obstacle courses that neither direct optimization nor hand-designed curricula could, and that movement between problems turned out to be essential Can paired environment and agent optimization unlock unsolvable challenges?. SAGA goes a step further and evolves the objective itself. An outer loop proposes new goals and compiles them into scoring code, so the goal stops being a fixed human input Can agents evolve their own objectives during search?. This is where human design really drops out: not when agents talk to each other, but when the environment and the goal stop being fixed by people.
Does the result actually match human work? In at least one case, yes. An agent rewritten through seven automated revisions over eight days equaled or beat its human-engineered counterpart on four held-out benchmarks, including weather forecasting Does automated evolution match human-built agent performance?. Most of this progress happens in the cheap "fast loop" of prompts, memory, and tools, not in model weights, because scaffold changes are cheap and easy to undo Do self-improving agents really split into two distinct loops?. There's a catch, though. Models of every size are about equally good at proposing improvements to their harness, but mid-tier models benefit most from them. Weak models don't use the changes, and strong ones don't follow them faithfully Do stronger models always evolve harnesses better?. So co-evolution doesn't automatically pay off more as models get more capable.
The corpus doesn't directly measure how much human design effort any of these systems saves, so "reduces reliance" holds up in performance terms, not in a counted tally of human work. Humans still choose the starting setup, what gets shared, and the outer loop, at least until the third stage of the survey's ladder.
Sources 9 notes
A survey framework organizes co-evolving systems into three stages that progressively remove human engineering: dynamic peers first, then adaptive environments and feedback, finally the evolution mechanism itself. Single-entity self-improvement stalls in static contexts; co-evolution supplies adaptive pressure across multiple components.
Group-Evolving Agents outperformed isolated tree-based self-evolution by 14–20 percentage points by explicitly pooling code patches and execution traces within each generation. Analysis showed five of eight key tool improvements came from different parent agents, proving the sharing mechanism itself—not just more search—drove the gains.
SkillClaw aggregates interaction trajectories across users, processes them through an autonomous evolver that identifies patterns and refines skills, then synchronizes updates system-wide. This converts siloed individual learning into shared capability improvement without manual curation.
Sequence model agents trained against diverse co-players develop in-context best-response strategies that naturally resolve into cooperation. Mutual vulnerability to exploitation creates pressure that drives cooperative mutual adaptation without hardcoded assumptions or timescale separation.
POET co-evolves environments and agents while allowing solutions to transfer between problems, producing sophisticated behaviors and solving obstacle courses that direct optimization and curriculum learning cannot. Transfer between environments proved essential to the system's success.
Show all 9 sources
SAGA's bi-level architecture closes a feedback loop from optimization results back to goal design by having an outer LLM loop propose new objectives and compile them into code the inner loop can immediately use, enabling objective formulation as part of discovery rather than a fixed input.
AIDE85, evolved through seven accepted rewrites in 8 days, equals or surpasses AIDEhuman on four held-out benchmarks spanning in- and out-of-distribution tasks including weather forecasting. The result shows automated design iteration can match human-driven R&D on generalization.
A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.
Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- Self-Improvements in Modern Agentic Systems: A Survey
- Rethinking the Evaluation of Harness Evolution for Agents
- Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators