Can an agent rewrite its own modification algorithm during runtime?
Gödel Agent enables LLMs to modify not just their policy but also the algorithm that decides how to modify themselves. This raises questions about whether true self-reference and unconstrained design freedom are achievable in practice.
Gödel Agent proposes a three-tier taxonomy of agent design freedom and places itself at the top tier. "Hand-Designed Agent" systems follow "the same policy π all the time, regardless of environmental feedback"; "Meta-Learning Optimized Agent" systems update that policy "based on a meta-learning algorithm I... at training time," but the meta-algorithm I itself "remains unchanged once deployed." Gödel Agent instead lets the agent "freely decide its own routine, modules, and even the way to update them," updating both the policy and the updater together each step: (πt, It) becomes (πt+1, It+1). The paper calls this self-reference — "the property of a system that can analyze and modify its own code, including the parts responsible for the analysis and modification processes" — and places it as the "highest degree of freedom" in its design-space figure, inspired by Schmidhuber's Gödel machine.
Mechanically, the agent's main loop is a recursive function rather than a fixed loop, in which an LLM decision function chooses among four actions: self_inspect (read its own current code from runtime memory), interact (call the environment's utility function U to score the current policy), self_update (have an LLM rewrite (π, I)), and continue_improve (recurse if no other action applies). The implementation achieves self-awareness and self-modification through "monkey patching" — reading and overwriting Python's own runtime memory — so the agent rewrites its running code during the loop that is rewriting it, guided "solely by high-level objectives through prompting," with no predefined routine or fixed optimization algorithm constraining the search.
This is a concrete, code-level instance of the first dial in What separates self-improvement from policy improvement?: Gödel Agent not only puts the improver I inside the agent, it makes I itself a target of self_update, so the dial's inside/outside boundary collapses into one self-modifiable object. The standard it is judged against, the utility function U, stays external and fixed — by the second dial this reads as "anchored," despite the paper's own label "self-referential." Against Do self-improving agents really split into two distinct loops?, it sits in the fast, non-parametric loop — fine-tuning its own LLM modules is named future work, not current behavior — but extends that loop's scope from prompts, memory and tools to the update algorithm itself. Can AI systems improve themselves through trial and error? is the later, named successor in the same Gödel-machine lineage: where Gödel Agent runs one recursive self-modification trajectory, DGM replaces that single trajectory with an evolutionary archive of variants validated empirically on benchmarks.
The excerpt gives no figures comparable to DGM's measured SWE-bench and Polyglot jumps — only that results "demonstrate... significant performance gain... surpassing various widely-used agents," unquantified here. The paper's own Limitations and Discussion sections flag the rest: it does not compare against the most engineered existing systems (OpenDevin is named explicitly as out of scope); it does not find "the exact point at which the agent can no longer comprehend and improve itself" as self-generated complexity grows, citing Yampolskiy (2015); and its Safety Considerations section argues, rather than demonstrates, that "fully self-modifying agents will require human oversight and regulation" as capability increases. At this strength the excerpt establishes feasibility of one self-referential, code-level loop, not a measured advantage over harder baselines or an account of when self-reference breaks down.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What limits recursive self-improvement in autonomous AI systems?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
What separates self-improvement from policy improvement?
Does recursive self-improvement work by the same evaluate-and-improve cycle as classical policy iteration, or are they fundamentally different processes? Understanding this distinction matters for predicting which self-improving systems remain controllable.
Gödel Agent embodies the first dial literally: the improving algorithm I sits inside the agent and is itself updated alongside the policy
-
Can AI systems improve themselves through trial and error?
Explores whether replacing formal proof requirements with empirical benchmark testing enables AI systems to successfully modify and improve their own code iteratively, and what mechanisms prevent compounding failures.
DGM later replaces Gödel Agent's single-trajectory recursive update with an evolutionary archive of variants validated on benchmarks
-
Do self-improving agents really split into two distinct loops?
Explores whether modern self-improving agents can be understood through a clean abstraction separating fast scaffold updates from slow model weight updates, and whether this framework actually explains the field's recent progress.
Gödel Agent sits in the fast scaffold loop but extends it to the update algorithm itself, not just prompts or tools
-
Can agents learn from vague goals without predefined metrics?
Most self-improving AI systems optimize toward explicit objectives. But what if an agent must first decide what capability to build, how to build it, and how to measure progress—all from only a natural-language goal?
Gödel Agent is guided solely by high-level objectives, leaving what to improve open — exactly the coupling this note says existing work underaddresses
-
How does an AI agent improve its own research code?
Explores the feedback loop where an AI research agent modifies and tests its own codebase, with each successful change becoming the agent that proposes the next revision. This specificity matters because it distinguishes a narrow, defined mechanism from broader claims about open-ended self-improvement.
Qualifies scope: AIDE2 limits recursive self-improvement to harness-layer edits, narrower than Gödel Agent's full self-modification
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
- Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- The Darwin Gödel Machine: AI that improves itself by rewriting its own code
- Hyperagents
- A Self-Improving Coding Agent
- Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
Original note title
Gödel Agent is self-referential because it modifies both its policy and its own self-modification algorithm, achieving the highest degree of design freedom