SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

Do agent improvements discovered in one model transfer to others?

When an AI agent discovers ways to improve its own performance on one foundation model, do those improvements carry over if you swap in a different model underneath? This matters for understanding whether self-improvement finds genuinely useful design principles or just model-specific tricks.

Synthesis note · 2026-10-08 · sourced from Frontier AI Risk & RSI

Sakana AI and Jeff Clune's lab, announcing the Darwin Gödel Machine (DGM), report that the self-improvements their coding agent discovers are not narrow tricks tied to one setup but generalize across both the underlying foundation model and the task domain. "The improvements discovered by the DGM (e.g., better tools, refined workflows) generalize to produce higher performance across different underlying FMs," they write: "an agent optimized with Claude 3.5 Sonnet also showed improved performance when powered by o3-mini or Claude 3.7 Sonnet." The same held across task domain — a DGM variant whose self-improvement was "guided exclusively by its performance on Python tasks" went on to show "significant performance gains on tasks in entirely different programming languages (like Rust, C++, and Go)."

The lab frames this as evidence that DGM "discovers general agent design improvements rather than just model-specific tricks." The self-modification loop — propose a code change, evaluate it on SWE-bench or Polyglot, keep what scores higher — is selecting for changes to the agent's tools and workflow (better patch validation, file viewing, solution ranking, a record of past failed attempts), not for changes tuned to exploit one model's quirks or one language's syntax. Because the archive keeps many agent lineages running in parallel rather than hill-climbing a single trajectory, the lab argues this open-ended exploration is what lets DGM land on transferable design improvements instead of overfitting to the benchmark it was scored against.

This sits alongside Can AI systems improve themselves through trial and error?, which covers the same system's benchmark gains (SWE-bench 20%→50%, Polyglot 14.2%→30.7%) and its empirical-validation-over-formal-proof mechanism; this source's distinct addition is the transfer result — that the discovered improvements hold up when the frozen foundation model underneath is swapped out, or when the task moves to a language the agent never trained on. That fits the scaffold/weights split in Do self-improving agents really split into two distinct loops?: DGM only ever rewrites the scaffold (tools, workflow) while the foundation model stays frozen, and the transfer finding is what makes that scaffold-only path look valuable — the gains are properties of agent design, portable across whichever model sits underneath it. It also echoes, at a smaller scale, the same generality that Can machine feedback sustain discovery at test time? claims for AlphaEvolve's evolutionary-archive discoveries carrying over from one deployed problem to another.

The excerpt gives no numbers for the cross-model transfer — only that performance "also" improved under o3-mini and Claude 3.7 Sonnet, not by how much relative to the 20%→50% SWE-bench gain reported for the original setup — so the strength of the transfer, as distinct from its mere existence, is unmeasured here. The lab's safety framing (sandboxed execution, human supervision, "please see our paper for a more thorough review") is asserted rather than demonstrated in this post, and the claim that self-improvement "could offer a novel pathway to enhance AI safety" is left unelaborated. What the excerpt does support is narrower and still useful: a self-modifying coding agent, evaluated under the pressure of two different benchmarks, arrived at design changes that are properties of agent design rather than of one model or one language — grounds for treating scaffold-level self-improvement through evolutionary search as producing findings worth testing elsewhere, rather than assuming they are benchmark artifacts.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can smaller specialized models match frontier models on key metrics? How should systems validate code that agents generate? What limits recursive self-improvement in autonomous AI systems?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 87 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Sakana AI reports Darwin Gödel Machine's self-discovered improvements transfer across models and programming languages