INQUIRING LINE

Instead of judging an AI by how smart it looks today, what if you scored it by how well its future upgrades actually turn out?

How does clade-level metaproductivity compare to true optimal self-modification decisions?

This explores whether you can judge a self-modifying agent by how well its descendants (its 'clade' or lineage) end up performing, and how close that judgment gets to the theoretically ideal choice about which self-changes to keep.


This explores whether you can score a self-modifying agent by how well its descendants perform, rather than by how good it looks right now, and how close that comes to the theoretically ideal choice about which self-changes to keep. One caveat up front: the collection doesn't hold the work that introduced 'clade-level metaproductivity' (the Huxley-Gödel Machine line of research), so it can't settle the formal comparison. That research argues that if you knew a modification's true lineage payoff, you could make the same accept-or-reject choices as a Gödel Machine, the idealized agent that only rewrites itself when it can prove the rewrite helps. In practice the lineage payoff has to be estimated from a limited number of descendants that were actually tried. What the collection does offer is several angles on the main problem underneath: a change's immediate score and its long-run value are often different things.

The closest mechanical parallel comes from tree search. Can tree structure alone convert outcome rewards into process supervision? shows that a branching structure can turn one final outcome into credit for each step, by comparing how sibling subtrees turn out. Scoring a clade works the same way: a node is judged by what grows beneath it. Can tree search replace human feedback in LLM training? makes the same move for training data, where the tree's results stand in for human labels. Both teach the same lesson. Subtree estimates are only as good as the branches you can afford to explore, and that cost is the gap between an estimated lineage score and the true optimum.

There is a direct example of why scoring by immediate performance misleads. Do stronger models always evolve harnesses better? finds that models of every strength write roughly equally useful edits to their own scaffolding, but how much they gain from those edits peaks at mid-tier models. So the quality of a self-edit and its actual payoff come apart, and a selector that only looks at the edit itself will pick wrong. Lineage scoring is one attempt to measure the payoff instead.

The bigger caution is that any score of this kind is still an external measuring stick. Can models reliably improve themselves without external feedback? argues that self-improvement that works always brings in an outside anchor, and a clade's score is anchored to whatever benchmark the descendants are tested on. That fits Are self-refinement and recursive self-improvement actually the same thing?: what works today is bounded, checkable refinement, not open-ended self-rewriting. A practical detail from Do self-improving agents really split into two distinct loops? explains why lineage search is workable at all: most self-modification now happens in prompts, memory and tools rather than in model weights, so many branches can be tried cheaply and abandoned.

The takeaway you might not have expected: the hard part of self-improvement may be less about making good changes and more about knowing which changes will pay off several generations later. That makes it a credit-assignment problem, the same one that runs through tree search and reinforcement learning.


Sources 6 notes

Can tree structure alone convert outcome rewards into process supervision?

Tree-GRPO uses branching structure to transform trajectory-level outcome rewards into step-level preference signals through sibling subtree comparison, eliminating the need for separate process reward models or step-level annotation while scaling with computational budget.

Can tree search replace human feedback in LLM training?

AlphaLLM uses tree search outcomes and three critic models to derive dense reward signals equivalent to human-labeled feedback. Tree structure naturally ranks solution paths by success, replacing the annotation oracle that standard RLHF requires.

Do stronger models always evolve harnesses better?

Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

Are self-refinement and recursive self-improvement actually the same thing?

A 1,250-paper survey shows that bounded, evaluable self-refinement (current industrial practice) differs fundamentally from open-ended recursive self-improvement, which remains constrained by grounding requirements, collapse dynamics, and compute limits measurable today.

Show all 6 sources
Do self-improving agents really split into two distinct loops?

A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.