A Self-Improving Coding Agent
Recent advancements in Large Language Models (LLMs) have spurred interest in deploying LLM agents to undertake tasks in the world. LLMs are often deployed in agent systems: code that orchestrates LLM calls and provides them with tools. We demonstrate that an agent system, equipped with basic coding tools, can autonomously edit itself, and thereby improve its performance on benchmark tasks. We find performance gains from 17% to 53% on a random subset of SWE Bench Verified, with additional performance gains on LiveCodeBench, as well as synthetically generated agent benchmarks. Our work represents an advancement in the automated and open-ended design of agentic systems, and demonstrates a data-efficient, non gradient-based learning mechanism driven by LLM reflection and code updates.
Introduction. LLMs have recently made impressive advancements across a range of domains and tasks [2, 14, 29]. However, in order to put these LLMs to use in real world applications, LLMs must be wrapped in code to orchestrate them and expose tools that allow the models to take actions. These action-taking LLMs are referred to as agents, and the broader system an agent system.
These agent systems often show dramatic improvements in benchmark performance over “plain” LLMs [43, 46, 5], through combinations of prompting strategies and methods for combining different LLM outputs. Early examples include best-of-N sampling and simple prompting strategies such as chain of thought [20]. However more sophisticated schemes have shown success in getting the desired behavior and performance improvements from the models, for instance STaR [45], Tree of Thoughts [42], Graph of Thoughts [3], LLM Debate [10], Iterative Self-Refinement [23], Expert Prompting [22] among many others. The comprehensive survey of Schulhoff et al. [34] demonstrates the vast number of manually created strategies to date.
Recent improvements in coding agents raise the question of whether these agents themselves can autonomously modify and improve their own code by discovering e.g. new prompting schemes or tools without manual design and implementation. We argue that this style of fully self-referential meta-agent programming is possible today and offers a sound alternative to the ad-hoc, trial-and-error approach of hand-crafted orchestrators, which may only explore a small fraction of the solution space. Recent work in the Automated Design of Agentic Systems (ADAS) [16] uses a meta-agent to optimise agent implementations. However, Hu et al. [16] is not self-improving, as there are two separate agents: the target-agent that performs the task, and the meta-agent, which improves the target agent. A motivation for a self-improving system is that the improvements in coding abilities may be leveraged during subsequent improvement steps, hopefully compounding.
Our contributions are:
• A self-improving coding agent (SICA) that eliminates the distinction between meta-agent and target agent, and is capable of editing its own codebase to improve itself with respect to its cost, speed and benchmark performance. • Empirical evidence that self-referential agents can effectively improve their own implementations; we find performance improves from 17% to 53% performance on a random subset of SWE-Bench Verified, even with consideration given to safety constraints and resource efficiency. • We share a our implementation of a self-improving coding agent (SICA) with the community. SICA is implemented in standard Python without a domain-specific language, and provides a reference agent framework for building new SICA systems, as well as those seeking to post-train LLMs on tool use and other agent tasks.
Related work. The traditional approach to developing and optimizing agent systems has been to manually design agent architectures and prompting techniques. Notable examples include Chain-of-Thought prompting [39], self-refinement [23] and self-reflection [37] for improving reasoning, tool use frameworks [33], and various compositional agent systems [38, 1]. While these hand-crafted approaches have achieved strong results, they require significant human effort and may miss useful patterns that could be discovered through automated search.
Another direction focuses on enabling agents to learn reusable skills and continuously self-improve. MaestroMotif [19] uses LLM feedback to learn skill rewards and combines skills through code generation. This builds on earlier work on intrinsically motivated reinforcement learning [6] and autotelic agents [8] that develop repertoires of internally motivated skills, as well as work in openendedness [48, 12] that use LLMs to identify interesting and useful directions to explore in.
Perhaps the most closely related line of prior work began with ADAS [16]. ADAS used a target-agent which performs the actual task, and a meta-agent which improves the target-agent. As such, ADAS is not self-improving (as the meta-agent improves the target agent, not itself). Moreover, in ADAS the meta-agent edits only a single forward function, written in a domain-specific language which has been carefully designed to make expressing different prompting schemes very straightforward. In contrast, our self-improving coding agent is fully self-improving (i.e. there is no distinction between the meta and target agent), and it operates over the agent’s full Python codebase.
Of course, we would expect the first truly self-improving agents to be coding agents, because agents are written in code. The natural approach is to start off with a basic coding agent that can open/close/edit files, run commands in the terminal etc, then to launch this agent in a selfimprovement loop. We believe that our self-improving coding agent is the first such work. However, there are two papers claiming self-improving agents, but they do not evaluate in the coding setting, as they do not consider “full” coding agents. First, Gödel Agent [44] has specific tools (such as action_adjust_logic and action_read_logic) that allow modification of small parts of the agent as it is running. Thus, it is not a general-purpose coding agent, as it is traditionally understood. And as such, as with ADAS, it was evaluated on language understanding and mathematical benchmarks (DROP [11], MGSM [36], MMLU [15] and GPQA [32]), rather than coding benchmarks. Second, Zelikman et al. [47] introduce a self-taught optimizer for recursively self-improving code generation.
Method. The main running loop of SICA resembles the ADAS loop Hu et al. [16] and is illustrated in Figure 1. In particular, both SICA, and ADAS keep an archive of previous agents and their benchmark results. In the first step, SICA takes the best performing agent from the archive up until this point as the meta-agent, which is instructed to look through the archive, identify an improvement, and implement it. The high-level scheme is illustrated in Algorithm 1. Note that this is different from ADAS, in that ADAS has a fixed meta-agent, so does not select the meta-agent from the archive (the archive in ADAS is an archive of target-agents, not meta-agents). Second, both ADAS and SICA evaluate the new agent on a set of evaluation benchmarks and store the results.
We define the ‘best’ performing agent by a utility function, which incorporates common desiderata for a good agent; the normalised benchmark performance score pscore ∈[0, 1], the wall-clock time in seconds ptime, as well as the dollar cost pcost. The base utility is calculated as We begin by describing the initial coding agent. We then go on to describe the benchmark runner framework, and how this framework naturally allows us to create a self-referential task (that of improving the coding agent).
We were interested in useful, practically relevant coding agents capable of self-improvement. To that end, we started with a sophisticated coding agent, with a number of tools, sub-agents and oversight mechanisms. Of course, the tools, sub-agents and oversight mechanisms were all amenable to self-improvement; but here we simply describe the starting point.
We designed a set of tools to open (i.e. add to the context) and close files, overwrite files, execute a shell command and a simple calculator tool. We also include basic tools to submit a final answer to a benchmark problem, return a result from a sub-agent call to its caller, as well as exiting early from a sub-agent call. Finally, we provide an ‘archive analysis’ tool, which facilitates extracting high-level summary statistics about each agent iteration in the archive. We describe the tool calling mechanism in more detail in Appendix C.
Next, we included several sub-agents. These are merely agents invoked by an existing agent, much like a tool call, by accepting input arguments and returning a string value. Introducing this abstraction serves two purposes. The first is to ensure the benchmark tasks can run entirely autonomously, without human direction and guidance. We provide the problem statement to the main entrypoint agent, (the full prompt for which is given in Appendix A) which then decomposes the task, delegating and asking questions to sub-agents. The second is to manage the context length for extended agent runs. The meta-improvement task is challenging, and the main agent may take on the order of one hour to finish one meta-improvement iteration. The subagent calls allow the context used for a sub-task to be naturally discarded when the agent returns, with the value returned by the sub-agent providing a summary of that sub-task in the main agent’s context. Initially, we included three agents: a coding agent, a problem-solver agent and a reasoning agent. The full prompts for these agents are given in Appendix A.1.
Finally, one important component we provide in the initial agent that enables extended, autonomous operation while remaining on-task, avoiding loops or other behaviours that frustrate progress is the asynchronous overseer. This is an LLM, running periodically in a concurrent thread, which is prompted to judge whether there are any pathological behaviours or deviations from the original task, and which may intervene by notifying the currently running agent (or any of its parents) by means of a message in-lined into the assistant context, or in serious cases to cancel the execution of an agent if it has deviated too far or appears stuck. The overseer is called every 30s by default, or on key events.
This overseer is provided with a comprehensive textual representation of the state of the system (a truncated example of which is provided in Appendix B) with verbose content truncated to limit context.
Discussion. It is vitally important to carefully consider the safety implications posed by systems designed to be self-improving. While our self-improving coding agent shows improvement in certain agentic tasks, there seem to be limits to what can be achieved with such purely scaffolding based systems (as seen in Figure 4), which do not modify the model weights. If the system updated the model weights as well as the scaffolding, we would expect that the risks of recursively self-improving systems would be higher. We highlight two key mechanisms for mitigating any risks.
First, a key safety mitigation in such systems is observability: that humans can and do oversee all aspects of potentially self-improving systems. We have thus architected SICA with sophisticated yet user-friendly observability mechanisms, which provide rich insight into the chain-of-thought, actions and sub-agent invocations at each step. These include an interactive web interface for easy oversight of all events occurring in the agent system, as well as the asynchronous LLM-based overseer, which may be prompted to treat various behaviors as grounds for canceling an agent, and which may also use a different LLM to the running agent. See Appendix A.2 for an example overseer judgement and the control afforded to it at each asynchronous invocation. Given the lack of model weight updates, which make it difficult but not impossible for the model to encode e.g. stenographic messages [24, 25] in the chain-of-thought, we believe this provides strong safety mitigation. That said, for recursively self-improving systems where the model weights are also updated, observability mechanisms, on their own, are clearly inadequate, and further research on safety mechanisms is clearly necessary.
The second opportunity for safety mitigation is inherent to the iterative, evaluation-based update procedure. While our initial benchmark set focuses on measuring capabilities we wish to improve, it is entirely possible to include safety-related evaluations in this benchmark set, to validate each agent before it progresses to the next iteration as the meta-agent.
We do not believe there to be any significant, direct adverse societal consequences to this work. Our objectives either focus on improving the mechanics of code editing or the effectiveness of multi-step reasoning through longer-horizon coding tasks.
Conclusion.
Limitations. Our initial attempt at a self-improving coding agent is not without limitations. One key difficulty was getting the LLM-based agent to autonomously come up with truely novel, innovative, feasible and interesting modification ideas at each meta-improvement step, which is a theme which has been commented on in the open-ended learning literature [27, 40, 17]. The cost of settling on a bad idea which suffered from poor ‘taste’ was a lengthy agent editing step followed by an even more expensive run through the benchmarks. While the failed iteration persists in the archive, in principle acting as an example of what not to do, we found that the initial feature ideas would often heavily influence later feature ideas as variations on the same theme. This path dependency may lead to higher variance agent runs; with poor quality initial feature suggestions (e.g. fixating on caching open files) often lowering the quality of subsequent feature suggestions.
We also note that in optimizing for agent running time and cost, our relatively short 5-minute timeouts (and to a lesser degree per-problem cost limits) cause the initial benchmark scores to perhaps be lower than expected for the underlying language model (e.g. Sonnet 3.5 v2), especially for longer-horizon benchmark tasks like SWE-Bench.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What limits recursive self-improvement in autonomous AI systems? Do individually safe AI actions create unsafe outcomes in integrated systems? Should governance of agentic AI systems be runtime or design-time?- What separates a self-sovereign agent from a merely rogue or misaligned one?
- How do controllable simulators compare to population-level agent simulation approaches?
- Can parallel agents or complementary mechanisms replace single-human interrogation of LLMs?
- Can single-model internal dialogue replace multi-agent debate systems?
- What makes the prompt a fundamentally new kind of speech act?
- Can prompting inject new knowledge into already-trained AI models?
- Can prompt engineering fully prevent role flipping in LLM agents?
- Why do LLM regenerations produce meaningfully different personalities from the same prompt?
- What distinguishes a neutral simulator from an agent with its own agency?
- How does the dialogue prompt establish the character the model plays?
- Can designated leadership structures reduce premature convergence in multi-agent reasoning?
- Why do multi-agent systems converge on wrong answers without debate safeguards?
- How does scene-switching prevent cross-problem interference in multi-agent reasoning?
- Can silent agreement be prevented in multi-agent reasoning systems?