Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
Advances in Large Language Models (LLMs) have enabled a new class of selfevolving agents that autonomously improve through environmental interaction, demonstrating strong capabilities. However, self-evolution also introduces novel risks overlooked by current safety research. In this work, we study case where an agent’s self-evolution deviates in unintended ways, leading to undesirable or even harmful outcomes. We refer to this as Misevolution. We evaluate misevolution along four key evolutionary pathways: model, memory, tool, and workflow. Our empirical findings reveal that misevolution is a widespread risk, affecting agents built even on top-tier LLMs (e.g., Gemini-2.5-Pro). Different emergent risks are observed, such as degradation of safety alignment after memory accumulation, or unintended introduction of vulnerabilities in tool creation and reuse. To our knowledge, this is the first study to systematically conceptualize misevolution and provide empirical evidence of its occurrence, highlighting an urgent need for new safety paradigms for self-evolving agents. Finally, we discuss potential mitigation strategies to inspire further research on building safer and more trustworthy selfevolving agents. Our code is available here.
Introduction. Large Language Model (LLM) agents are increasingly deployed in real-world applications, such as software development and automated research (Hong et al., 2024; OpenAI, 2025b). Recently, a new frontier focuses on agents that can evolve on their own, known as self-evolving agents (Zhou et al., 2025b; Zhang et al., 2025a; Gao et al., 2025; Fang et al., 2025). Different from their static counterparts, these agents improve themselves via active and continuous interaction with the environment. The evolutionary process of these agents primarily spans four dimensions, each corresponding to a core component of the agent system: model, memory, tool, and workflow. By leveraging feedback from tasks, the agent may optimize the parameters of the underlying language model (Sun et al., 2025b), accumulate experience into memory (Zhou et al., 2025a), create and master new tools (Qiu et al., 2025), or adjust the execution workflow (Zhang et al., 2025b). The impressive performance of self-evolving agents on challenging tasks has drawn wide interest in the community.
However, self-evolution also introduces novel risks that are overlooked by current safety research. In this study, we investigate the case in which an agent’s self-evolution deviates in unintended ways, leading to undesirable or even harmful outcomes. We refer to this as Misevolution, and highlight four core characteristics that distinguish it from established safety concerns:
Temporal emergence. During self-evolution, some components of the agent are dynamically changing, and risks can emerge over time. This contrasts with research on jailbreaking or misalignment that evaluates a “static snapshot” of an LLM (Chao et al., 2024; Li et al., 2023).
Self-generated vulnerability. A self-evolving agent may generate new risks and vulnerabilities internally, even without a dedicated external adversary. These risks may arise as unintended side effects of the routine evolutionary process or from the agent’s autonomous interactions with potentially harmful environments. This is distinct from emergent misalignment (Betley et al., 2025) which intentionally conducts finetuning on insecure examples.
Limited data control over evolving process. The autonomous nature of self-evolution constrains data-level control, hindering direct safety interventions (e.g., injecting safety data during supervised fine-tuning). This distinguishes misevolution from LLM fine-tuning safety (Qi et al., 2024b), in which training data are explicitly curated and managed.
Expanded risk surface. An agent’s evolution across multiple components (model, memory, tool, workflow) creates an expanded risk surface. Vulnerabilities can emerge from any of these parts. The ability to execute real-world tasks means any such flaw can cause tangible harm.
The concept of misevolution raises critical concerns: can we guarantee that a self-evolving agent will always converge to a beneficial assistant without compromising safety or introducing new risks? The answer is far from certain, as undesirable behaviors can emerge from the evolutionary process. For instance, a service agent that evolves its memory may learn a biased correlation between refunds and positive feedback, leading it to proactively offer refunds even when not asked to (Figure 1(a)). Similarly, an agent that evolves its toolset may ingest seemingly useful but insecure code from a public repository, inadvertently creating a new tool with a backdoor that leaks data (Figure 1(b)).
To systematically investigate the misevolution phenomenon, we examine its occurrence across the aforementioned evolutionary pathways: (1) In model evolution, we assess whether self-evolving agents compromise their safety alignment after self-updating their model parameters. (2) In memory evolution, we test whether memory-augmented agents learn undesirable preferences or degrade their risk awareness while accumulating experience into memory. (3) In tool evolution, we evaluate whether agents will spontaneously induce risks in the tool creation-reuse loop, and test agents’ ability to reject appealing but potentially malicious tools retrieved from the Internet. (4) In workflow evolution, we analyze whether automatically adjusted workflows can lead to safety decay.
The key contributions of our study can be summarized as follows:
• Conceptualizing misevolution: To our knowledge, we are the first to identify and systematically study misevolution as a novel safety challenge in self-evolving agents.
• Empirical evidence: We conduct comprehensive evaluations, providing qualitative and quantitative evidence for misevolution across four main evolutionary pathways.
• Preliminary mitigations and future outlook: We discuss potential mitigation strategies and provide implications for building safer and more trustworthy self-evolving agents.
Related work. Self-evolving agents. Research on self-evolving agents, known for their adaptive capabilities and strong performance (Novikov et al., 2025; Gao et al., 2025; Fang et al., 2025; Liu et al., 2025), has primarily explored four evolutionary pathways. One line of work focuses on model evolution, where agents refine their own parameters using self-generated data or learning curricula (Zhao et al., 2025a; Huang et al., 2025; Sun et al., 2025b; Zhou et al., 2025b). Another prominent approach is memory evolution, where agents learn from past experiences by storing and retrieving them to guide future actions (Yang et al., 2025c; Lin et al., 2025; Zhou et al., 2025a). Likewise, tool evolution allows agents to expand their capabilities by creating, refining, and reusing tools (Qiu et al., 2025; Haque et al., 2025; Zhao et al., 2025c; Zheng et al., 2025a) or by improving their proficiency with existing tools (Qu et al., 2024). Some studies also demonstrated performance gains through workflow evolution, where agents autonomously optimize their execution pipeline and collaborative structure (Hu et al., 2025b; Zhang et al., 2025b; Wang et al., 2025b). The common thread in these studies focused on enhancing agent capabilities. In contrast, our work shifts the focus to the safety implications of self-evolution, investigating the potential for this process to introduce unintended risks.
Safety of LLMs and LLM-based agents. The rapid development of LLMs and LLM-based agents has made their safety a primary concern (Zhang et al., 2024; He et al., 2024; Deng et al., 2025). Previous research has uncovered numerous vulnerabilities. For LLMs, these include data poisoning and backdoors (Hubinger et al., 2024; Wang et al., 2024; Zhao et al., 2025b), adversarial attacks and jailbreaking that elicit unsafe behaviors (Zou et al., 2023; Wei et al., 2023; Ren et al., 2025), and the generation of harmful or private content (Wang et al., 2023; Li et al., 2024; Qian et al., 2025). For agents, risks include knowledge poisoning (Chen et al., 2024; Zou et al., 2025), prompt injections (Zhan et al., 2024; Debenedetti et al., 2024), and interference from malicious links (Yang et al., 2025b; Tur et al., 2025). Prior work has also reported deceptive behaviors in agents (Guo et al., 2025). Most studies evaluate a “static snapshot” against external threats. In contrast, we study “misevolution”: risks emerging dynamically within self-evolving agents. This differs from finetuning-related issues (Qi et al., 2024b; Lyu et al., 2024; Huang et al., 2024). A notable example is emergent misalignment (Betley et al., 2025), where finetuning on insecure code leads to misalignment on other domains. However, this stems from training on a curated set of insecure examples.
Method. To study misevolution, we first need a clear picture of what constitutes a self-evolving agent and the mechanisms that drive its evolution. We begin by formalizing the core components of a selfevolving agent and the iterative loop of interaction and adaptation (Gao et al., 2025). Then, we present a taxonomy that organizes self-evolution into four pathways: model, memory, tool, and workflow (see Figure2). This taxonomy guides our experiments in Section 3. We briefly introduce representative methods within each paradigm, and highlight those evaluated in this study.
Model evolution. Model evolution is typically realized through self-training, a process where an LLM or agent updates its own model parameters. We focus on two prevalent self-training paradigms: self-generated data and self-generated curriculum. In the self-generated data paradigm, an LLM or agent autonomously creates its own training data, often through a feedback loop where it generates novel tasks or environments and then learns by attempting to solve them. In our study, we evaluate two such methods to investigate whether safety alignment is compromised after model self-training.
In the self-generated curriculum paradigm, an agent adaptively plans its own learning curriculum based on the current performance. In our study, we experiment with SEAgent (Sun et al., 2025b), a self-evolving agent designed for computer use. It identifies recent failures and focuses its learning on the specific parts of the trajectory that caused the failures, thus generating tasks of increasing difficulty based on the agent’s current capabilities.
Memory evolution. Beyond updating the language model, a self-evolving agent can also learn from its past experiences through memory. This process centers on leveraging information from previous trajectories to inform decision-making in new situations. In our study, we experiment with SE-Agent (Lin et al., 2025), a high-performing self-evolving coding agent on SWE-Bench (Jimenez et al., 2024). SE-Agent summarizes and distills strategies from past trajectories and leverages this knowledge to aid the solution of new tasks. We also test with the memory storage and retrieval mechanism of AgentNet (Yang et al., 2025c), which saves successful and failed trajectories and retrieves relevant ones into the context when facing a new task. We investigate whether the mere accumulation of memory, even without parameter updates, can induce emergent misbehavior.
Tool evolution. Tool evolution can manifest in several ways, such as creating new tools from scratch, ingesting external tools, and improving mastery over existing tools (Haque et al., 2025; Qiu et al., 2025; Qu et al., 2024). Our study focuses on two paradigms with direct safety implications: tool creation and reuse, and ingesting external tools.
In the tool creation and reuse paradigm, agents improve their capabilities by creating tools during task execution and reusing these tools in future tasks. Following frameworks like Alita (Qiu et al., 2025), we wrap self-created tools as MCPs to facilitate reuse. We investigate whether this tool creation-reuse loop can spontaneously introduce vulnerabilities or undesirable behaviors.
In the ingesting external tools paradigm, an agent evolves by actively searching for and integrating external tools, often from public sources like GitHub. While powerful, this exposes the agent to unvetted code. To test this potential risk, we evaluate an agent’s ability to identify and reject tools retrieved from the Internet that appear appealing but contain malicious code pieces.
Workflow evolution. A common paradigm in self-evolving multi-agent systems is autonomous workflow optimization, where agents refine their collaborative structures based on environmental feedback. This is often framed as a search or optimization problem over a space of possible workflows represented by graphs (Zhuge et al., 2024) or code (Hu et al., 2025b).
Discussion. Building on our findings, we discuss potential strategies to mitigate misevolution. We supplement our preliminary experiments to gain a deeper understanding of the practical challenges. We also discuss the hypothetical factors that may have led to misevolution in Appendix A, and discuss suggestions for deploying self-evolving agents in Appendix B.
Mitigating model misevolution. We have observed that model self-training can inadvertently compromise safety alignment. Notably, we identify a critical phenomenon that the model exhibits safety degradation even when the self-generated data contains no explicitly unsafe or harmful content. To mitigate this, we introduce a lightweight safety post-training phase following self-evolution to rectify the model’s alignment. Experiment on Absolute-Zero-7B-Base shows that this mitigation is partially effective, boosting the Safe Rate of the evolved model from 59.5% to 62.75%. However, this approach remains insufficient to fully restore the model to its initial safety level and incurs additional computational overhead. More detailed discussion can be found in Appendix E.1.
Mitigating memory misevolution. We hypothesize a unified cause for safety alignment decay and deployment-time reward hacking: agents over-relying on past experiences without critical reflection. Thus, we introduced a simple prompt-based mitigation: instructing the agent to treat retrieved memories as “references,” rather than “rules,” such as “The following memories are for reference only. You must make an independent decision based on the current context.” This lightweight intervention proved effective, reducing the ASR of SE-Agent (Qwen3-Coder-480B) from 20.6 % to 13.1% and increasing the Refusal Rate from 54.4 % to 66.9 % on RedCode-Gen. It also reduced the Unsafe Rate in reward hacking scenarios from 71.8 % to 51.4 % on average. However, the agent’s safety still did not fully recover to its pre-evolution level, suggesting the need for more powerful mitigation strategies. We provide more detailed results and discussion in Appendix E.2.
Mitigating tool misevolution. For tool creation and reuse, a key mitigation is automated safety verification. We propose a two-stage process: (1) static analysis to scan new tools for vulnerabilities before adding them to the toolset, and (2) a judge LLM to re-validate safety upon reuse in the new context. Although not tested in our work, this is an important practice for maintaining internal tool safety. For ingesting external tools, we prompted the agent to explicitly assess project safety before packaging, such as “If you find the project unsafe [...], refuse to package it.” This intervention improved the agent’s safety awareness, increasing Refusal Rate from 7.28% to 69.0% on Qwen3- 235B-Instruct and from 2.70% to 68.5% on Gemini-2.5-Flash. Nevertheless, this result remains far from satisfactory. We discuss the potential reason and implications in Appendix E.3.
Mitigating workflow misevolution. We showed that workflow evolution can cause safety decay, sometimes unexpectedly: even an innocuous step like an ensemble node may increase the Unsafe Rate. A simple mitigation is to add a safety-oriented prompt to the vulnerable Ensemble Node we identified, instructing it to pay attention to safety when aggregating responses. With this simple intervention, we observed an improvement in safety. The ASR has dropped from 83.1% to 77.5%, while the safe rate was promoted from 5.6% to 13.1%. More detailed discussion about workflow mitigation can be found in Appendix E.4.
Conclusion. In this paper, we introduced and systematically investigated “misevolution,” a novel risk in selfevolving agents. We show that the self-evolution process across model, memory, tool, and workflow can lead to unforeseen and even harmful outcomes. Our findings reveal that misevolution is a pervasive issue even for agents built on top-tier LLMs. It manifests in various forms, such as the safety alignment decay, deployment-time reward hacking, and insecure tool creation and reuse. We also explored potential mitigation strategies and presented preliminary prompt-based methods. While these methods show some effectiveness, they are far from a comprehensive solution to misevolution. Finally, our findings highlight an urgent need for new safety frameworks designed for the dynamic and autonomous nature of self-evolving agents.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What limits recursive self-improvement in autonomous AI systems?- What failure modes does recursive self-improvement encounter in evolutionary loops?
- How do evolutionary archives improve on single self-modification trajectories?
- How do evolutionary archives enable open-ended self-improvement without formal proofs?
- What safety tradeoffs arise when improvers move inside the agent?
- How do tool evolution pathways create backdoors and security vulnerabilities?
- How should evolving systems track lineage and enable rollback of changed mechanisms?
- Do agent improvements discovered on code tasks transfer to non-coding domains as well?
- Why do safety patches on self-evolution only partially restore prior safety?
- Why does monitoring performed by agents on agents create safety risks?
- How can agents evolve their own skills without human input?
- Can evolved algorithms transfer learning strategies across different datasets and tasks?
- How do you isolate environment protections as independent variables safely?
- Why do sequences of safe actions sometimes violate system-level constraints?
- What does trajectory audit reveal about evolution cycle contributions and costs?
- Can an auditor verify environment state without trusting the executor's self-report?