INQUIRING LINE

Why does training an AI on an easy problem sometimes unlock a hard one it couldn't solve directly — or even step by step?

How do stepping stone solutions transfer effectively between different environments?

This explores how a solution worked out for one task or environment can become a useful starting point for a different one, and what makes that kind of hand-off work rather than fail.


This explores how a solution that works in one setting becomes a starting point, or stepping stone, for another, and why that sometimes beats attacking the hard problem directly. The clearest case in the collection is POET, a system that invents new obstacle courses and trains agents on them at the same time. It also lets an agent trained on one course be tried on the others Can paired environment and agent optimization unlock unsolvable challenges?. The surprising part is that this transfer was what made the system work. Some courses that couldn't be solved by training on them directly, or by a hand-designed easy-to-hard curriculum, were solved by an agent that had picked up its skills somewhere else. The route to a hard problem often goes through a problem nobody would have chosen as the next step.

Why does a transferred solution help rather than hurt? The best clue comes from Sakana AI's Darwin Gödel Machine. It is an agent that rewrites its own tools and workflows, and it keeps an archive of earlier versions. Improvements found with one model still worked with different models (Claude 3.5 Sonnet, o3-mini, Claude 3.7 Sonnet). Improvements found on Python tasks also helped in Rust, C++ and Go Do agent improvements discovered in one model transfer to others?. The likely reason is that the improvements fixed general agent design, not quirks of one model. That points to a rule of thumb: a stepping stone carries over when it captures something about the shared structure of the problems, not a trick tuned to one setting. A study of reward hacking defenses makes a similar split. Some defenses work the same way across training setups, while others only have loose analogues in other setups Which reward hacking defenses actually transfer across training substrates?.

The same idea shows up under other names. Training web agents with a curriculum that gradually allows longer interactions works as a series of stepping stones, each one building on the last Does agent interaction time scale separately from reasoning depth?. Distillation research shows the same thing between models. A student learns more reliably from a teacher that stays close to it than from a strong teacher far ahead Can proximity between teacher and student fix distillation instability?. On-policy distillation, where a teacher grades the student's own attempts, mostly helps the student find good paths it could already reach. It doesn't raise the student's ceiling Does on-policy distillation actually expand student capability?. Together these suggest a limit: a stepping stone is useful when it sits within reach of where the learner already is.

A caveat: the collection shows that stepping-stone transfer works, mainly through POET and the Darwin Gödel Machine. It has much less on how to tell in advance which solutions will transfer, or how to pick which environment to borrow from. If that's the question you care about, these two systems are the place to start.


Sources 6 notes

Can paired environment and agent optimization unlock unsolvable challenges?

POET co-evolves environments and agents while allowing solutions to transfer between problems, producing sophisticated behaviors and solving obstacle courses that direct optimization and curriculum learning cannot. Transfer between environments proved essential to the system's success.

Do agent improvements discovered in one model transfer to others?

Sakana AI's Darwin Gödel Machine discovered improvements to agent tools and workflows that transferred to different foundation models (Claude 3.5 Sonnet, o3-mini, Claude 3.7 Sonnet) and to programming languages outside its training domain (Python-trained agents improved on Rust, C++, Go), suggesting the improvements target portable agent design rather than model-specific exploits.

Which reward hacking defenses actually transfer across training substrates?

A systematic map identifies which defense mechanisms function identically across weights, selection, and text substrates versus which only provide functional analogies. Practitioners rated this correspondence analysis as their most immediately useful takeaway from the work.

Does agent interaction time scale separately from reasoning depth?

Test-time interaction—increasing environment steps—enables exploration, backtracking, and replanning that per-step reasoning cannot achieve. Curriculum-based RL on rollout length produces SOTA web agents, showing interaction scaling dominates on tasks with partial observability.

Can proximity between teacher and student fix distillation instability?

TOP-D constructs a close teacher instead of distilling from a distant target, bounded by a trust region. This controls gradient variance, guarantees monotonic improvement, and outperforms standard distillation with zero computational overhead.

Show all 6 sources
Does on-policy distillation actually expand student capability?

On-policy distillation steers students toward correct reasoning paths within their existing capability envelope rather than raising the ceiling. Signal quality and diversity matter far more than teacher scale; a smaller teacher with high-fidelity guidance outperforms larger teachers without it.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.