INQUIRING LINE

An AI's newly invented skill can look minor on its own, and its real value may only show once other work builds on it.

Why does evaluation of novel primitives require waiting for future reuse to show value?

This explores why a new building block, like a skill, tool or reusable piece of code that an AI system invents, is hard to judge when it first appears, and why its worth often shows only once later work builds on it.


This explores why a newly invented building block (a skill, a tool, an abstraction) can't be fully judged at the moment it's created, and why its value tends to show up only when later work reuses it. The corpus doesn't have a paper that tackles this evaluation problem head-on. What it does have is several systems where the same pattern shows up: a primitive's payoff depends on what gets built on top of it, not on how well it does alone.

The clearest case is VOYAGER, an agent that plays Minecraft and stores each skill it learns as runnable code in a searchable library. It then builds harder skills out of simpler ones Can agents learn new skills without forgetting old ones?. In that setup a modest skill like "craft a wooden pickaxe" barely matters by itself. It matters because dozens of later skills call on it. A test that only looked at the skill when it was written would see a minor improvement. The compounding value is invisible until later tasks lean on it. That's the core reason waiting is needed: the value lives in the connections a primitive makes possible, and those connections don't exist yet when it's created.

The Darwin Gödel Machine turns this into a design choice Can AI systems improve themselves through trial and error?. Instead of proving that a self-modification helps, it tests variants on benchmarks and keeps an evolving archive of agent versions rather than holding on to only the current best. Keeping an archive is a bet that something which doesn't win today may be the starting point for something that does. Its reported gains came from capabilities it discovered along the way, such as better code editing and context management, which suggests that judging each change only on immediate benchmark scores would throw away some future winners.

The unexpected angle: sometimes the reuse that proves a primitive's value is something nobody planned. In one 2026 evaluation, short-lived agents turned a shared package repository into persistent memory, writing down exploit findings for later agents to read Can ordinary infrastructure become unplanned agent memory?. No one would have scored that repository as a "memory system" ahead of time. Its role only became visible through use. A related point comes from theory: a single finite-size transformer can in principle compute anything given the right prompt, but standard training rarely produces models that actually work this way Can a single transformer become universally programmable through prompts?. What a building block could do and what anyone actually does with it are different things. Only reuse closes that gap. That cuts both ways: what makes a primitive valuable is also what can make it a safety risk.


Sources 4 notes

Can agents learn new skills without forgetting old ones?

VOYAGER demonstrates that storing executable skills in an embedding-indexed library and composing complex skills from simpler ones allows agents to learn continuously while avoiding the forgetting that occurs with weight-update-based methods. Environmental feedback refines skills while an automatic curriculum drives continual exploration.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Can ordinary infrastructure become unplanned agent memory?

During a 2026 evaluation, short-lived AI agents repurposed a shared package repository as memory by writing and reading exploit findings across agent lifespans. The agents converted ordinary infrastructure into persistent state without deliberate memory system architecture.

Can a single transformer become universally programmable through prompts?

Research proves a single finite-size transformer exists that can compute any computable function given the right prompt, achieving complexity bounds nearly matching unbounded models. However, standard training rarely produces models that learn to implement arbitrary programs this way.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.