Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
Introduction. Open-endedness and cumulative progress are key characteristics of scientific breakthroughs [1, 2, 3]. However, most existing AI systems rely on pre-defined model architectures designed by humans. Although such systems can accumulate experience through training, they often struggle to transcend the capability boundaries imposed by their initial designs, as they lack the ability to modify their own structural configurations [4]. Thus, progress remains heavily dependent on continuous human intervention.
Existing open-ended self-improving systems are largely inspired by biological evolution and designed around individual-centric evolutionary processes [2, 4, 5, 6, 7]. At each iteration, a single agent is selected as the parent and refined to produce one or more offspring(Figure 1a). The overall structure follows chainor tree-structured evolution, where different branches remain strictly isolated. Consequently, although such systems often exhibit substantial exploratory diversity, this diversity rarely serves as effective stepping stones [8, 9]. Instead, many agents provide only temporary diversity, producing short-lived variants that fail to contribute to long-term cumulative progress.
It is time to rethink agent evolution. AI agents are not biological individuals; why should their evolution remain constrained by biological paradigms? In fact, AI agents can directly share trajectories, tools, and learned artifacts, and they can aggregate complementary skills without the constraints of reproduction or lineage.
Therefore, we introduce Group-Evolving Agents (GEA), a new paradigm for open-ended self-improvement that treats a group of agents, rather than an individual agent, as the fundamental unit of evolution (Figure 1b). This shift enables explicit experience sharing and reuse across agents within a group, naturally allowing exploratory discoveries from different agents to be consolidated and accumulated into long-term progress rather than remaining as short-lived variants. At each iteration, GEA first selects a parent group of agents using a Performance-Novelty criterion that balances immediate performance gains with evolutionary diversity. The parent agents then jointly produce a child group through a shared pool of aggregated experience from all members.
We evaluate GEA on challenging coding benchmarks, achieving success rates of 71.0% on SWE-bench Verified and 88.3% on Polyglot, significantly outperforming state-of-the-art open-ended self-evolving methods (56.7% and 68.3%, respectively). Analysis reveals that GEA more effectively consolidates the diversity generated during open-ended exploration, yielding sustained progress and stronger performance given the same number of evolved agents. By leveraging experience from better-performing agents, GEA also exhibits stronger robustness to framework-level perturbations. Furthermore, its improvements stem from workflow and tool enhancements rather than model-specific optimizations, thus transferring consistently across GPTand Claude-series models.
Additionally, by leveraging meta-learning for self-improvement in open-ended exploration, without any human intervention, GEA achieves performance comparable to or even surpassing human-designed state-ofthe-art frameworks on both benchmarks (71.0% vs. 71.8% on SWE-bench Verified, 88.3% vs. 52.0% on Polyglot).
In summary, we propose Group-Evolving Agents, a new paradigm for open-ended self-improvement that:
- Overcomes the limitation of inefficient utilization of exploratory diversity caused by branch isolation in existing tree-structured evolution, by enabling explicit experience sharing and reuse within the group during evolution. 2. More effectively consolidates and reuses experience and evolutionary diversity from other agents, achieving significant performance gains and stronger robustness over state-of-the-art open-ended self-evolving methods, with improvements that transfer consistently across different coding models. 3. Matches or surpasses human-designed state-of-the-art frameworks through meta-learning-based selfimprovement without human intervention.
Related work. Recent years have witnessed growing interest in how AI systems can continuously improve themselves without human intervention [10, 11, 12]. Most existing self-improving approaches mainly focus on continuous, iterative refinement of the given agent system [13, 14, 11, 15, 16], typically evolving toward a specific optimization objective and following a linear, chain-based evolutionary structure [17, 12, 6] . Such systems achieve self-improvement through mechanisms such as self-play against historical versions or self-generated verification [18, 19, 20, 21, 22], supervised fine-tuning [23, 24, 25] or reinforcement learning on selectively filtered feedback [26, 27, 28] , and reflection-based methods [29, 4, 13] or in-context learning [30, 31]. While this goal-oriented, chain-based evolutionary paradigm enables autonomous improvement along a particular direction, it inherently limits the ability of self-evolving systems to explore diverse evolutionary directions in open-ended solution spaces.
A line of work has pointed out that one of the key challenges in enabling unbounded improvement and innovation lies in developing open-ended AI systems that can continuously produce both novel and learnable artifacts [1, 2, 3, 32]. Building on this insight, open-endedness has been characterized as the capability of systems to continuously generate artifacts that are novel, interesting, and learnable from a human perspective [2, 33, 34, 35, 36, 37].
Motivated by the potential of enabling unbounded evolution through open-ended exploration in selfevolving agents, more recent studies adopt lineage-based, tree-structured evolutionary strategies [38] inspired by biological inheritance and mutation [2, 7, 39, 40]. In these frameworks, individual parent agents are selected at each iteration to independently produce offspring, enabling various branching exploration across multiple evolutionary directions and helping avoid local optima. However, the strict isolation between evolutionary branches prevents effective information and experience sharing and reuse across lineages. As a result, many promising directions discovered early in evolution persist only as temporary diversity and fail to contribute to long-term cumulative progress. To overcome this limitation, we introduce a group-centric evolutionary paradigm, Group-Evolving Agents (GEA), which explicitly enables intra-group experience sharing and reuse throughout the evolutionary process. By consolidating complementary discoveries across agents, GEA more effectively leverages the diversity generated by open-ended exploration to support sustained cumulative progress.
Method. We propose Group-Evolving Agents, a framework for open-ended evolution that treats a group of agents as the fundamental unit of evolution. GEA maintains an archive that stores all discovered agents throughout the evolutionary process. As shown in Figure 1, at each iteration, GEA proceeds in two core stages:
(1) Parent Group Selection (§3.1): GEA first selects K parent agents from the archive using a Performance– Novelty selection strategy [8, 9, 41] that balances immediate task-solving competence with long-term evolutionary diversity and potential.
(2) Open-ended Group Evolution (§3.2): The selected agents form a parent group that jointly produces an offspring group of the same size through explicit experience sharing and reuse across parent agents.
We detail the method below.
3.1 Parent Group Selection Inspired by Mouret and Clune [8], Pugh et al. [9], Chatzilygeroudis et al. [41], parent group selection in GEA balances two key principles: performance and novelty. We prioritize agents with strong task performance, as performance reflects an agent’s immediate competence and its likelihood of producing effective offspring, since evolution in GEA proceeds through iterative modifications of the agent’s implementation, which itself constitutes a form of solving coding problems. At the same time, we also encourage exploration beyond currently well-optimized regions of the search space, as agents that exhibit novel evolutionary directions may contribute to long-term cumulative progress even when their current performance is not optimal.
We represent each agent i using a task-success vector zi ∈{0, 1}D, where each dimension indicates whether the agent successfully solves a corresponding probe task. Similar binary task–response representations of this form have been widely used to characterize an agent’s coding capabilities and to better understand how these capabilities are distributed across various tasks [42, 43]. Using this representation, we measure the dissimilarity between two agents via cosine distance: where NM(i) denotes the set of M agents with the smallest cosine distance to agent i.
To construct the parent group, we rank agents according to a combined score nov(i) moderates the influence of novelty. Finally, we select the top-K agents according to this score to form the parent group. Performance serves as the primary selection criterion, while novelty is incorporated as a mild bias without dominating performance, enabling a balanced trade-off between exploitation and exploration. The full procedure is summarized in Algorithm 1.
3.2 Open-Ended Group Evolution Unlike conventional approaches where parent agents evolve independently without information and experience exchange, GEA explicitly enables experience sharing and reuse among agents during evolution. This group-level experience sharing allows agents to integrate complementary evolutionary directions explored by different agents while maintaining open-ended exploration. Diversity generated during exploration is thus transformed from transient variations into long-term useful experience, effectively contributing to sustained evolutionary progress.
Given a selected parent group G = {a1, a2, . . . , aK}, GEA generates a new group G′ of the same size, where each agent evolves by leveraging both its own evolutionary history and experience aggregated from other members of the parent group, as demonstrated in Figure 2.
For each agent ai ∈G, we collect a set of evolutionary traces consisting of:
- the code modification patches applied to the agent’s framework; 2. a predicted task patch generated by ai for a randomly sampled unsolved task during evaluation; 3. the corresponding task execution logs, including the complete tool invocation history and execution workflow; 4. the evaluation outcome of the same task, which exposes failure modes and potential directions for framework-level improvement.
Discussion. Overall, our analysis shows that GEA can efficiently consolidate tool-level innovations discovered across the agents, rather than letting them remain isolated in separate evolutionary branches. Figure 4 summarizes nine key tool-level modifications on agents’ framework that drove improvements. GEA integrated eight of these functionalities into its best agent, whereas the best DGM agent integrated only five. Crucially, the four tools missing from the DGM agent were explored in isolated branches (e.g., T4 at iteration 9) but failed to propagate due to lineage isolation. In contrast, GEA systematically consolidated these dispersed capabilities; Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing five of its integrated tools originated from different parent agents, confirming that explicit experience sharing prevents beneficial innovations from dying out.
Polyglot benchmarks. This indicates that the improvements induced by group-evolving persist across different backbone models.
Further analysis reveals that all performanceimproving patches discovered during GEA evolution, including those from the best agent and the top-3 performing agents, primarily target the agent’s workflow and tool usage rather than model-specific prompting, details can be found in Table 3 in Appendix.
These findings together with Figures 5 demonstrate that although GEA leverages a specific backbone model to drive evolution, it discovers agent-level improvements that are largely model-agnostic and the evolved agents could generalize across different coding models. experiences from better-performing agents to guide the repair of faulty ones, confirming the robustness of the group-evolving paradigm.
GEA demonstrates the potential and viability of group-evolving open-ended systems to autonomously modify their own implementation for continuous improvement. While this potential aligns with the goal of building AI that benefits humanity, open-ended exploration also carries inherent considerations worth noting. For instance, the evolutionary process may inadvertently introduce directions misaligned with human intent while consuming substantial computational resources, or produce patches that lack structural clarity, leading to increasingly complex systems that are difficult to fully understand. Therefore, it is essential to establish appropriate boundaries and guide the system to preserve exploratory diversity while ensuring alignment with human intent. Following Zhang et al. [2], all experiments in this work are conducted in isolated sandbox environments, thereby limiting potential impacts on host systems.
On the other hand, although we focus on evolving agents’ coding capabilities in this work, this paradigm has broader potential applications, for example, enabling systems to mitigate biases through self-improvement, thereby becoming more trustworthy and beneficial for social good.
Conclusion. We introduce Group-Evolving Agents (GEA), a new paradigm for open-ended self-improvement that treats a group of agents, rather than an individual agent, as the fundamental unit of evolution. By enabling explicit experience sharing and reuse within the group, agents can learn from each other’s evolutionary experiences and adaptively integrate complementary improvements throughout evolution.
Compared to individual-centric self-evolving approaches, GEA more effectively consolidates valuable exploratory outcomes from early stages into the best-performing agents, efficiently transforming transient diversity into long-term useful experience. As a result, group-level evolution achieves substantially stronger performance given the same number of evolved agents.
Further analysis shows that GEA’s improvements primarily stem from enhancements to agent workflows and tool usage, rather than overfitting to a specific coding model. Therefore, its gains transfer consistently across different models, including both GPT-series and Claude-series.
In addition, GEA exhibits stronger robustness than individual-centric self-evolving approaches: through group-level experience reuse, better-performing agents can guide the repair of faulty ones, enabling GEA to recover from framework-level bugs with fewer evolution iterations.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI agents improve their skills through accumulated experience and reuse? What limits recursive self-improvement in autonomous AI systems?- How do evolutionary archives improve on single self-modification trajectories?
- How do evolutionary archives enable open-ended self-improvement without formal proofs?
- What collapse dynamics constrain recursive self-improvement in current evidence?
- Do evolutionary discovery systems like FunSearch count as bounded or open-ended improvement?
- Does co-evolution empirically outperform single-entity self-improvement in standard evaluations?
- Why does the generation-verification gap limit what an agent can improve about itself?
- Do evolutionary archives let agents improve themselves without formal proof?
- How did individual agents shift toward collective swarm behavior?
- Does co-evolution between peer agents reduce reliance on human design?
- Can pluralism survive within a single platform or does it require architectural exits?
- Can agents cooperate through self-modeling when incentive structures are fundamentally misaligned?
- Do agent improvements discovered on code tasks transfer to non-coding domains as well?
- How do tool evolution pathways create backdoors and security vulnerabilities?