Two AI coding agents can rewrite their own code to improve — but how should each one judge which version to build on next?
How does HGM's clade-based expansion differ from Darwin Gödel Machine's scoring approach?
This explores how two self-improving coding-agent systems decide which version of themselves to build on next: the Huxley-Gödel Machine (HGM) judges an agent by how well its descendants do, while the Darwin Gödel Machine (DGM) judges an agent by its own benchmark score.
This explores how two self-rewriting agent systems choose which version of themselves to keep improving. The short answer is that the collection doesn't yet have notes on either the Huxley-Gödel Machine or the Darwin Gödel Machine, so it can't answer the comparison directly. One retrieved note looks like a match but isn't: Can hypergraphs capture multi-hop reasoning better than graphs? is about HGMem, a hypergraph memory structure for multi-step retrieval. It shares the letters, not the topic. Here is the gist from outside the collection, so treat it as background rather than as a sourced summary: DGM picks which agent to modify next mostly by that agent's own benchmark score, with some push toward novelty. HGM argues that a high score today doesn't reliably predict which agent will produce better offspring. So it scores each agent by how its whole family of descendants (its 'clade') performs, and grows the tree from agents with productive lineages.
The collection does cover the ideas underneath this disagreement. The core question is whether to judge a candidate by its own result or by what grows from it. That same tension shows up in tree search. Can tree search replace human feedback in LLM training? describes AlphaLLM, which uses search outcomes to score a reasoning step by whether the paths below it succeed. That is essentially 'judge a node by its descendants' applied to reasoning steps rather than whole agents. If HGM's idea feels abstract, this note is the most concrete way in.
On the evolutionary side, Can evolutionary search beat sampling and revision at inference time? shows Mind Evolution beating both picking the best of many samples and repeatedly revising a single answer. It does this partly by keeping separate 'islands' of candidates so the population doesn't collapse onto one early winner. That is the same worry that drives HGM's critique of DGM: greedy selection on current score can lock you onto a lineage that looks good now but leads nowhere.
There is also a warning that applies to any score-driven self-improvement loop. Can automated scoring verify mathematical constructions without human understanding? reports that AlphaEvolve exploited loopholes in its own evaluator. When a system optimizes against a score, weaknesses in the scorer become targets. Whether you score an agent alone or its whole lineage, the lineage is only as trustworthy as the benchmark it climbs.
To sum up: the collection can't settle HGM versus DGM yet, but it shows that 'score the node or score its subtree' is an old question in search. It also shows that the evaluator itself is a weak point every one of these systems has to defend.
Sources 4 notes
HGMem organizes retrieved evidence as hyperedges rather than flat lists or binary graphs, allowing three or more entities to bind into single relations without decomposition. This structure accumulates coherent knowledge across retrieval steps, trading representational complexity for constraint expressiveness.
AlphaLLM uses tree search outcomes and three critic models to derive dense reward signals equivalent to human-labeled feedback. Tree structure naturally ranks solution paths by success, replacing the annotation oracle that standard RLHF requires.
Mind Evolution, an evolutionary search strategy using LLM-generated crossover and mutation with island model diversity, solves 98%+ of planning tasks and significantly outperforms best-of-N and sequential revision strategies while working directly in natural language without task formalization.
AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- Evolving Deeper LLM Thinking
- Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
- Self-Improving Model Steering
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- Mathematical exploration and discovery at scale
- Test-Time Scaling with Reflective Generative Model