Can paired environment and agent optimization unlock unsolvable challenges?
Does co-evolving environmental challenges alongside agent solutions, with transfer between problems, enable systems to solve obstacles that neither direct optimization nor standard curricula can crack?
POET (Paired Open-Ended Trailblazer) "pairs the generation of environmental challenges and the optimization of agents to solve those challenges," running many paths through the space of problems and solutions at once, and critically "allows these stepping-stone solutions to transfer between problems if better, catalyzing innovation." Tested in a 2D bipedal-walking obstacle-course domain, POET "produces a diverse range of sophisticated behaviors that solve a wide range of environmental challenges, many of which cannot be solved by direct optimization alone, or even through a direct-path curriculum-building control algorithm introduced to highlight the critical role of open-endedness." The paper states plainly that "the ability to transfer solutions from one environment to another proves essential to unlocking the full potential of the system as a whole."
Mechanically, POET maintains a list of environment-agent pairs and runs three operations each iteration: it generates new environments by mutating an active environment's encoding, admitting the mutation only if the originating agent showed enough progress to make reproduction worthwhile and if the new environment is "neither too hard nor too easy for the current population" (priority goes to the most novel candidates); it optimizes each paired agent against its own environment (via evolution strategies in the experiments, maximizing whatever performance measure the environment poses); and it attempts to transfer each agent's current network to other active environments, keeping the transfer if it outperforms the resident agent. The active-environment population is capped and pruned oldest-first, giving agents time to optimize and their skills time to transfer before an environment is retired. This generalizes the minimal-criterion idea from MCC and the niche-optimization of quality-diversity algorithms, adding explicit coevolution between problems and solutions plus cross-environment "goal switching" as the transfer mechanism.
Against Can AI systems invent new concepts rather than reuse trained ones?, POET sits entirely on the search side of that later paper's distinction: it mutates a fixed, bounded genome rather than inventing new representational primitives, so it never confronts the vocabulary gap. Its minimal-criterion check is a cheap, fast verifier precisely because the representation stays fixed — the kind of "fast, cheap, decisive" evaluation that paper says only holds inside a fixed frame, not the verifier gap it describes for judging a genuinely new primitive. Against Can agents learn new skills without forgetting old ones?, POET's transfer step is the coevolutionary analogue of VOYAGER's skill library: instead of retrieving skills by embedding similarity, POET empirically tests each agent against every active environment and keeps whichever performs best, so stepping stones are found by trial rather than by semantic retrieval. The lineage also runs forward to Can AI systems improve themselves through trial and error?, which keeps POET's archive-of-stepping-stones structure but applies it to an agent rewriting its own code and validating empirically, rather than to a population of coevolving environment-agent pairs.
The excerpt tests one domain (2D obstacle courses) with a bounded genome: the paper's own limitations section concedes the environment space can "max out," since "there is a maximum possible gap width and stump height." Open-endedness here means diversification within a fixed, pre-specified generative space, not the invention of new problem representations — the paper's language about "indefinite" or billion-year-scale open-endedness is aspirational framing, not a demonstrated result. What is measured is narrower and still notable: coevolutionary generation plus empirical transfer solves configurations that direct optimization or a direct curriculum, run on the same fixed encoding, do not. The implication holds only within that scope — this is an existence proof for stepping-stone transfer inside a bounded search space, not evidence that the same coevolutionary loop scales to expanding the representation itself.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can smaller specialized models match frontier models on key metrics? Does AI-assisted research sacrifice exploration breadth for productivity gains? When do multi-agent systems improve over single frontier models?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can AI systems invent new concepts rather than reuse trained ones?
Current AI systems excel at search and reasoning within fixed representational frames, but can they autonomously create novel primitives like mathematicians invented negative numbers or entropy? This matters because genuine open-ended innovation may require frame-altering operations, not just frame-internal search.
contrast: POET mutates a fixed encoding and never confronts the vocabulary gap that paper identifies as the harder problem
-
Can agents learn new skills without forgetting old ones?
Explores whether externalized skill libraries—storing learned behaviors as retrievable code rather than parameter updates—can solve the catastrophic forgetting problem that plagues continual learning systems.
parallel mechanism: POET's cross-environment transfer is the coevolutionary analogue of VOYAGER's embedding-retrieved skill library
-
Can AI systems improve themselves through trial and error?
Explores whether replacing formal proof requirements with empirical benchmark testing enables AI systems to successfully modify and improve their own code iteratively, and what mechanisms prevent compounding failures.
lineage: DGM keeps POET's stepping-stone archive structure but applies it to an agent rewriting its own code
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
- Artifacts as Memory Beyond the Agent Boundary
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
- Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
- Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
Original note title
POET pairs environment generation with agent optimization so transferred solutions become stepping stones to otherwise unsolvable challenges