Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine

Paper · arXiv 2510.21614 · Published October 24, 2025
Frontier AI Risk & RSI

Recent studies operationalize self-improvement through coding agents that edit their own codebases. They grow a tree of self-modifications through expansion strategies that favor higher software engineering benchmark performance, assuming that this implies more promising subsequent self-modifications. However, we identify a mismatch between the agent’s self-improvement potential (metaproductivity) and its coding benchmark performance, namely the Metaproductivity- Performance Mismatch. Inspired by Huxley’s concept of clade, we propose a metric (CMP) that aggregates the benchmark performances of the descendants of an agent as an indicator of its potential for self-improvement. We show that, in our self-improving coding agent development setting, access to the true CMP is sufficient to simulate how the Gödel Machine would behave under certain assumptions. We introduce the Huxley-Gödel Machine (HGM), which, by estimating CMP and using it as guidance, searches the tree of self-modifications. On SWEbench Verified and Polyglot, HGM outperforms prior self-improving coding agent development methods while using fewer allocated CPU hours. Last but not least, HGM demonstrates strong transfer to other coding datasets and LLMs. The agent optimized by HGM on SWE-bench Verified with GPT-5-mini and evaluated on SWE-bench Lite with GPT-5 achieves human-level performance, matching the best officially checked results of human-engineered coding agents. Our code is publicly available at https://github.com/metauto-ai/HGM.

Introduction. Processes of self-modification drive the growth of complex systems, from biological evolution (Hendrikse et al., 2007; Dawkins, 2019) to cultural and scientific innovation (Good, 1966; Hall, 2007). These general ideas have been instantiated in concrete algorithms for self-improving agents (Schmidhuber, 1987; 2003; Nivel et al., 2013; Everitt et al., 2016), demonstrating how abstract principles of self-modification can be translated into operational mechanisms. Unlike static systems constrained by fixed architectures, such agents can incrementally modify their own selfmodification mechanisms and learning strategies, reusing newly gained abilities to fuel subsequent improvements. This capacity fosters continual adaptation, reduces reliance on human intervention, and enables problem-solving capabilities that cannot be fully anticipated at design time.

A central challenge is how to decide which self-modifications to accept. The Gödel machine (Schmidhuber, 2003) (GM) offers a theoretically optimal answer: accept only modifications that provably increase the expected long-term utility. While this provides a sound blueprint, its reliance on formal proofs makes it practically challenging. Recent implementations instead rely on coding agents that edit their own codebases and favor self-modifications from agents with higher benchmark performance (Robeyns et al., 2025; Zhang et al., 2025a). Yet, as illustrated in Figure 1 (left), this heuristic can be misleading: a high-scoring agent may produce unproductive descendants, while a lower-scoring one seeds lineages that achieve greater long-term gains. We term this phenomenon the Metaproductivity–Performance Mismatch.

To address this mismatch, we introduce clade-level metaproductivity (CMP), inspired by Huxley’s notion of clades as lineages of common ancestry (Huxley, 1957). CMP quantifies the productivity of a clade by aggregating the success of an agent’s descendants rather than relying solely on its immediate benchmark score. Furthermore, we show in Theorem 1 that in our self-improving coding agent development setting (Assumption 1, which includes the assumption that the only quality of the self-improvement process is the evaluation score of the final agent and that the evaluation is conducted with repeatable trials), having access to the true CMP oracle suffices to imitate the Gödel Machine.

This insight motivates our proposed algorithm, the Huxley–Gödel Machine (HGM), which approximates GM-style self-improvement by estimating CMP from clade-aggregated descendant outcomes and selecting nodes to expand via Thompson sampling. Furthermore, by leveraging a more reliable estimate, we adaptively decouple expansion from evaluation, leading to asynchronous execution for efficient parallelism.

To summarize, our contributions are as follows:

• We analytically define the Clade-Metaproductivity (CMP) function as a measure of agents’ selfimproving ability and show that in a self-improving coding agent development setting (Assumption 1), access to a CMP oracle suffices to reproduce the Gödel Machine’s acceptance mechanism. (Theorem 1). • We empirically observe that immediate benchmark performance is an unreliable predictor of CMP and show that our CMP estimator aligns better. • Using our CMP estimator, we propose the Huxley–Gödel Machine (HGM), which approximates the Gödel Machine in a coding agent setting from partial evaluations and guides the expansion via Thompson sampling with adaptive scheduling. • We empirically validate HGM on SWE-bench Verified and Polyglot, demonstrating higher-quality optimized agents compared to previous self-improving methods, even though they were discovered within substantially smaller allocated CPU-hours. Furthermore, HGM achieves human-level coding agent design on SWE-bench Lite by optimizing on SWE-bench Verified.

Related work. The general concepts of machine self-improvement were first systematically articulated by Good (1966), who described the possibility of “Intelligence Explosion" once machines acquire the capacity to design more capable successors. Early work on explicit self-improvements dates back to Schmidhuber (1987), which introduced self-referential learning mechanisms in which a system generates and evaluates modified descendant versions of itself. Follow-up work on selfimprovement progressed through interaction and agentic reinforcement learning. The Success-Story Algorithm(SSA) (Schmidhuber & Zhao, 1996; Schmidhuber et al., 1997) progressively forces selfmodifying policies to discover more effective self-modification strategies. Its core mechanism is based on hindsight: at each checkpoint, a sequence of self-modifications that did not yield higher long-term reward rates is systematically undone. In this way, SSA enforces continual improvement by ensuring that only those self-modifications associated with demonstrably greater reward intake per unit time are preserved. Fitness-Monotonic Execution (Kirsch & Schmidhuber, 2022a;b) reduces the outer-loop design by favoring the execution of models with higher ancestral performance. Meta-discovered update rules optimized optimizers (Metz et al., 2021) and black-box search (Lange et al., 2023). On the other hand, the Gödel Machine, a fully self-referential algorithm that rewrites its own code whenever it can prove an expected-utility improvement, provides a provably and globally optimal mechanism for self-improvement (Schmidhuber, 2003).

The rise of contemporary LLMs has created an opportunity to automate substantial aspects of software engineering. One concrete step in this direction is the development of coding agents, which extend LLMs with the ability to operate in conventional computing environments. ChatDev (Qian et al., 2023) first illustrated this idea in the context of automated bug fixing, and similar frameworks were later explored in SWE Yang et al. (2024), OpenHands (Wang et al., 2024), MetaGPT (Hong et al., 2024), and AgentLess (Xia et al., 2025).

The Self-Taught Optimizer (Zelikman et al., 2024) and Gödel Agent (Yin et al., 2024) first experimented with agents that modify their own scaffolding. Subsequently, DGM (Zhang et al., 2025a) and SICA (Robeyns et al., 2025) extend this direction by implementing self-modifying machines as full software engineering projects, where agents self-reference and modify their own repositories while validating changes through execution-grounded software engineering tasks. Both DGM and SICA, explicitly or implicitly, assume that higher software benchmark scores correspond to greater self-improvement capacity.

Method. Both the Darwin Gödel Machine (DGM) and the Self-Improving Coding Agent (SICA) belong to the class of self-referential AI (Schmidhuber, 1987; 2003). In particular, in DGM and SICA, agents modify themselves to generate new agents, each empirically validated on downstream tasks.

In this paper, we formalize this self-improvement process as an iterative tree-search problem, where the goal is to discover an agent that maximizes performance across multiple downstream tasks. Concretely, starting from an initial agent as the root, a tree-search policy incrementally grows the tree of self-modified agents. At each iteration, the policy either selects an agent (a node in the tree) to expand by producing a child agent (a self-modified version of the selected agent) or selects an agent to undergo additional evaluation on downstream tasks.

Formally, let Tt denote the archive of our agents at iteration t. In this paper, the archive is always represented as a tree of evolved agents, and we use the terms archives and trees interchangeably. T0 = {a0} is initialized as a single-node tree with a fixed initial agent. At each iteration t, the policy selects actions at+1 ∼π(· | Tt), where π is a policy over actions At = Mt ∪Vt, Mt = {ma : a ∈ Tt} are agent modifications, and Vt = {va : a ∈Tt} denotes evaluations. Action ma instructs the agent a to produce a self-modification that is added as a child of a to the tree, and va selects agent a from the tree for an additional evaluation on one more downstream task. After exhausting the computational budget, the policy selects a final agent (afinal = arg maxa∈T Scoreπ(a) ∈TB where B is the termination iteration and Score is part of the policy) from the final tree as the returned agent. The objective is to optimize J(π) = E[U(afinal)], where U is a utility function that measures the performance of downstream tasks. In this work, we consider U the average of binary success indicators across all downstream tasks. π denotes an algorithm, with DGM, SICA, and our proposed HGM representing concrete instances.

Compound Policy. At each step of self-improvement, the system faces a compound decision: whether to expand the tree by generating new agents or to evaluate existing ones. This decision naturally decomposes into three sub-policies: (i) a selection policy that chooses between expansion and evaluation, (ii) an expansion policy that determines which parent to modify, and (iii) an evaluation policy that selects an agent to test. Prior approaches, such as SICA and DGM, conflate these choices. They always expand a parent, create a child, and immediately evaluate that child on multiple tasks. This fixed sequence restricts flexibility: once a new agent is generated, it monopolizes evaluations, even if older agents appear more promising. For instance, an agent that fails nine tasks in a row continues to consume evaluations, while an older agent with partial successes is ignored.

HGM breaks this rigidity by decoupling expansion from evaluation. At each step, it adaptively decides whether to generate a new agent or to further probe an existing one, and evaluations are always at the granularity of a single agent–task pair. This finer control enables early stopping on unpromising agents. Table 5 summarizes how SICA, DGM, and HGM instantiate these sub-policies.

In this section, we introduce the Huxley–Gödel Machine (HGM), a self-improving machine that approximates Gödel Machine by using clade-level statistics. At the core of HGM lies the notion of metaproductivity—a measure of an agent’s ability to improve its self-improvement skills, which leads to better downstream performance of distant future agents.

The original Gödel Machine is a general task solver that, in principle, can optimally make any provable self-improvements in any computable environment with respect to a given objective (Schmidhuber, 2003). It achieves this by running a proof searcher, continually looking for formal proofs that some modification of its own code will yield higher expected utility. Once such a proof is found, the modification is executed and permanently alters the machine.

Conclusion. In this work, we identify a key limitation in the search heuristics of current self-improving coding agents: Benchmark scores alone do not reliably indicate an agent’s long-term potential for selfimprovement, since high-scoring agents can still lead to stagnating lineages, while seemingly weaker ones may seed productive self-improvements. We refer to it as the Metaproductivity–Performance Mismatch. To address this gap, we introduce Clade-Metaproductivity (CMP), a lineage-based metric inspired by Huxley’s notion of clades. We show that, under certain assumptions, when applied to our self-improving coding agent search problem (Assumption 1), the CMP oracle is sufficient to implement the Gödel Machine (Theorem 1).

Building on this principle, we propose the Huxley–Gödel Machine (HGM), which approximates CMP and uses it to guide expansion through Thompson sampling with adaptive scheduling. Empirically, HGM consistently produces higher quality agents than prior self-improving frameworks while also reducing wall-clock time. Notably, HGM generalizes across both dataset and model shifts, achieving human-level coding agent design performance on SWE-bench Lite with GPT-5 despite being optimized on SWE-bench Verified with GPT-5-mini.

Taken together, these results suggest that clade-based measures of improvement potential, rather than immediate performance alone, lead to more effective forms of self-improvement. By demonstrating that clade-level evaluation can reliably guide the growth of coding agents, this work points to a new paradigm for the design of agentic improvement: one in which improvement is driven not by narrow benchmarks, but by the long-term generative potential of entire lineages. This perspective underscores the importance of systems that strengthen an agent’s capacity to keep improving over time, rather than merely boosting their performance in the short term.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What human oversight must AI research systems have? What limits recursive self-improvement in autonomous AI systems? Do single-axis benchmarks accurately measure agent capability for real deployment? Why does AI verification capability persistently exceed generation capability? Can AI systems achieve real improvement without external human feedback? Why does polished AI output gain credibility despite fundamental verifiability problems? Why do confident AI outputs mislead human trust calibration? Should models ask for clarification when facing ambiguous or under-specified information? Why do standard evaluation practices obscure safety-critical AI failures? How should systems validate code that agents generate? Can AI systems discover fundamental improvements to their own architectures? How do models learn from self-generated outputs without cascading failures? When do multi-agent systems improve over single frontier models?