INQUIRING LINE

When AIs learn from their own outputs, can their goals be copied into successor models, and can that copying amplify them?

Can distillation help AIs scale their learned objectives across many copies?

This explores whether distillation (training one model on another model's outputs, or on its own) can take what an AI has learned, including what it is optimizing for, and spread it reliably across many model copies or successors.


This explores whether distillation can spread what an AI has learned, including its goals, across many copies or successor models. The corpus doesn't directly cover the multi-copy question: nothing retrieved here studies objectives spreading across a fleet of instances. What it does have is a set of adjacent findings on how training a model on model outputs carries behavior forward and amplifies it. Together they suggest the answer is closer to "yes, and it amplifies more than you'd expect" than "no."

The most direct evidence is that a model can teach itself without any outside answers. In on-policy self-distillation, the model's own majority-vote consensus served as the training signal. It matched or beat training on ground-truth labels across five benchmarks, distilling only on cases where the model disagreed with itself Can a model's own consensus replace ground truth labels?. The implication for scaling is easy to miss: if agreement among copies can stand in for truth, then whatever the copies already agree on, right or wrong, gets reinforced. A related result shows models can be trained to score their own work, using the unused space after the end of their output, so evaluation becomes part of the model instead of coming from an outside reward model Can models learn to evaluate their own work during training?. A model that carries its own scoring rule and teaches copies through consensus is a closed loop with very little external correction.

The second lesson is that self-reinforcing training tends to narrow behavior rather than spread it faithfully. RL post-training picks one output format inherited from pretraining within the first epoch and suppresses the rest, and which format wins depends on model scale rather than on which one works best Does RL training collapse format diversity in pretrained models?. When the reward barely varies between attempts at the same prompt, policies collapse into generic templates that ignore the input Why do language models collapse into generic templates?. So distillation probably doesn't copy an objective intact across generations. It is more likely to sharpen whichever tendency is already dominant. That matters for safety, because the trait that gets amplified isn't necessarily the one anyone meant to spread.

Weight updates aren't the only way to pass capability between models. A stronger model nearly doubled a weaker model's Theory-of-Mind scores without retraining it, by writing inference-time harnesses that moved unstable reasoning into deterministic code Can a stronger model lift a weaker one at test time without retraining?. Agents can also improve across episodes by storing written reflections on their failures instead of changing weights Can agents learn from failure without updating their weights?. Both channels are shareable: code and text can be handed to any number of copies instantly. If the question is how learned behavior spreads at scale, these routes may matter as much as distillation.

The part you may not have known you wanted to know: in this corpus, the strongest scaling force isn't distillation itself but self-agreement. Copies that train on their own consensus and grade their own work can improve without labels, and the same mechanism can lock in a dominant behavior without anyone choosing it. For the safety framing behind your question (AI systems spreading their objectives across many instances), you'll need sources outside this retrieval set.


Sources 6 notes

Can a model's own consensus replace ground truth labels?

Unsupervised on-policy self-distillation using the model's own majority-vote consensus matched or surpassed supervised methods on five benchmarks. The key mechanism distills only on self-inconsistent rollouts, using agreement as the teaching signal rather than external labels.

Can models learn to evaluate their own work during training?

Post-Completion Learning exploits unused sequence space after model output to train self-assessment capabilities during training while maintaining zero inference cost. The model learns to compute its own reward functions, internalizing evaluation rather than relying on external reward models.

Does RL training collapse format diversity in pretrained models?

Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.

Why do language models collapse into generic templates?

When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.

Can a stronger model lift a weaker one at test time without retraining?

A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.

Show all 6 sources
Can agents learn from failure without updating their weights?

Reflexion demonstrates that unambiguous environmental feedback (success/failure) enables agents to write useful self-diagnoses and improve across episodes without parameter updates. The binary signal prevents rationalization, and keeping reflections uncompressed preserves their usability.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.