INQUIRING LINE

Why does keeping a whole gallery of past attempts around find better ideas than endlessly polishing just one?

Why do evolutionary archives lead to more transferable discoveries than single-trajectory optimization?

This explores why AI systems that keep a growing library of past attempts, and let ideas move between them, tend to find solutions that work beyond the problem they were found on, compared with systems that keep refining one solution step by step.


This explores why keeping a population or archive of past attempts, instead of polishing a single solution, tends to produce discoveries that carry over to new problems. The corpus supports one part of this strongly and another only partly. It shows clearly that archives beat single-path search. It shows less directly that what they find transfers better. The mechanism behind the first part, though, explains why the second is plausible.

The clearest evidence is that good ideas often come from somewhere other than the path you're on. In Group-Evolving Agents, pooling code changes and execution logs across a whole generation beat isolated lineages by 14–20 points. Five of the eight key tool improvements came from a different parent than the agent that ended up using them Does sharing experience across agents beat isolated evolution?. POET makes the same point with environments: agents trained on one obstacle course were moved onto others, and that transfer was essential to solving courses that direct optimization never could Can paired environment and agent optimization unlock unsolvable challenges?. A single trajectory can only build on its own history, while an archive lets a solution found for problem A become the stepping stone for problem B. Being reusable across contexts is built into the search itself.

Diversity is the second ingredient. Mind Evolution keeps separate 'islands' of candidates and recombines them, and it beats both best-of-N sampling and step-by-step revision of a single answer Can evolutionary search beat sampling and revision at inference time?. Frontis-MA1 makes recombination something a model can learn: it trains a model on the moves of evolutionary search, including a 'Crossover' step that merges two programs, and the training and the search add to each other's gains Can training and search gains add together in program evolution?. The intuition is that a solution that survives being combined with others is probably capturing something general, not a trick that only fits one setup. AutoML-Zero hints at this from the other direction. Evolving from bare math operations, it rediscovered broadly useful techniques like weight averaging and learning-rate decay, not narrow hacks Can evolutionary search discover machine learning algorithms from scratch?.

An archive is also an asset in its own right. Dream-RSI replays its accumulated tree of discoveries as a cheap simulator for testing new exploration strategies, so the history keeps paying off after the original search ends Can past discoveries train better exploration policies?. SkillClaw does the same across people: it pools interaction histories from many users so that a skill learned in one person's context becomes everyone's How can agent systems share learned skills across users?.

The honest gap: the most direct evidence of transfer comes from AIDE2. Its gains held on four held-out benchmarks, including weather forecasting, which was outside its training distribution Do AIDE2's improvements transfer to unseen tasks?. But no paper in the retrieved set compares an archive against a single path and measures held-out transfer side by side. The surprising takeaway is about where the advantage comes from. It doesn't come mainly from searching more. It comes from letting solutions travel between problems and lineages, and that only works in domains with fast, clear scores to evaluate against What makes a research domain suitable for autonomous optimization?.


Sources 9 notes

Does sharing experience across agents beat isolated evolution?

Group-Evolving Agents outperformed isolated tree-based self-evolution by 14–20 percentage points by explicitly pooling code patches and execution traces within each generation. Analysis showed five of eight key tool improvements came from different parent agents, proving the sharing mechanism itself—not just more search—drove the gains.

Can paired environment and agent optimization unlock unsolvable challenges?

POET co-evolves environments and agents while allowing solutions to transfer between problems, producing sophisticated behaviors and solving obstacle courses that direct optimization and curriculum learning cannot. Transfer between environments proved essential to the system's success.

Can evolutionary search beat sampling and revision at inference time?

Mind Evolution, an evolutionary search strategy using LLM-generated crossover and mutation with island model diversity, solves 98%+ of planning tasks and significantly outperforms best-of-N and sequential revision strategies while working directly in natural language without task formalization.

Can training and search gains add together in program evolution?

Frontis-MA1 trained a single model on four program-evolution operators (Draft, Improve, Debug, Crossover), then reused those operators in long-horizon evolutionary search. The result was complementary gains: learning and search improved performance together rather than substituting for each other, lifting Medal Average from 39.39% to 71.21% on MLE-Bench Lite.

Can evolutionary search discover machine learning algorithms from scratch?

AutoML-Zero evolved algorithms from 65 basic operations that match neural networks and rediscover modern techniques like weight averaging and learning-rate decay, adapting strategies to task conditions in controlled experiments.

Show all 9 sources
Can past discoveries train better exploration policies?

Dream-RSI demonstrates that accumulated discovery trees can be replayed off-policy to score exploration policies without repeated online evaluation. The framework loops between policy evaluation on historical data, online redeployment, and simulator expansion, reportedly achieving competitive discovery quality at lower cost.

How can agent systems share learned skills across users?

SkillClaw aggregates interaction trajectories across users, processes them through an autonomous evolver that identifies patterns and refines skills, then synchronizes updates system-wide. This converts siloed individual learning into shared capability improvement without manual curation.

Do AIDE2's improvements transfer to unseen tasks?

The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.

What makes a research domain suitable for autonomous optimization?

Autonomous research pipelines require immediate scalar metrics, modular architecture, fast iteration cycles, and version control. Domains lacking any property resist autoresearch regardless of LLM capability, because the bottleneck is environmental structure, not model power.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.