Line of inquiry
Inquiring lines›What drives capability improvement…›How do training signals and method…›this line of inquiry
How do models learn from self-generated outputs without cascading failures?
A broader line of inquiry — a family of 54 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 54
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why does self-correction during generation produce reliable labels without exemplars?
- How does error avalanching compound failures in self-training iterations?
- How does error distribution during training affect a model's ability to self-correct?
- Can models learn to generate their own training examples effectively?
- Can self-consistency checks fully prevent error avalanching in self-training loops?
- What failure modes emerge when model-generated content trains on itself iteratively?
- Why does self-consistency fail as a proxy reward for correctness?
- Why does self-generated training data outperform externally sourced data?
- Why does self-generated training data outperform externally curated domain examples?
- What role does the self-consistency threshold play in preventing error reinforcement?
- How does self-distillation differ from standard fine-tuning approaches?
- Does self-generated training data reduce a model's capability diversity?
- Can self-training drift be prevented by applying student compatibility filtering?
- Does the generation-verification gap define where self-rewarding actually works?
- What happens when models train on feedback from their own generations?
- Can AI-generated explanations of errors teach as effectively as self-resolution?
- What makes policy self-distillation more effective than external teacher distillation?
- Why do error avalanches accelerate in self-training loops without verification?
- How does adversarial collapse threaten unsupervised self-play skill construction?
- What makes self-consistency a sufficient training target for the judge role?
- Why does self-judgment of success or failure work without ground truth labels?
- What happens when models train on AI-generated content recursively?
- Why does pure self-improvement require an external signal to produce genuine gain?
- How does distilling only inconsistent rollouts compare to distilling all generations?
- Can synthetic self-play data teach models when to disagree?
- How do instruction backtranslation and MAGPIE demonstrate self-generation principles?
- Can self-ratings of output quality predict forecast performance?
- How does distribution mismatch between training and deployment break self-correction?
- Does self-play feedback improve skills created from the agent's own experience?
- Do correlated human errors prevent models from transcending their training sources?
- Does keeping original real data present prevent irreversible model collapse?
- Why does systematic overconfidence on self-generated outputs compound autoregressive errors?
- Why does unsupervised self-distillation lose its advantage in thinking model mode?
- Do self-generated explanations outperform passive instruction for oversight?
- Why does filtering for correct examples prevent error compounding in self-training?
- Does self-conditioning improve belief-behavior alignment better than external priors?
- What causes code quality to degrade across multiple rounds of recursive self-training?
- What causes irreversible model collapse when training on model-generated content?
- Can models detect statistical properties of their own generation in real time?
- Can unsupervised confidence-based training scale to domains beyond human evaluation reach?
- How does adversarial self-play during training differ from single-model self-revision?
- Can the serving loop itself become the primary training data source?
- Does model collapse happen identically with unfiltered versus filtered synthetic data?
- Can self-distillation reduce catastrophic forgetting in continual learning?
- Why does online RL succeed where supervised training fails for self-correction?
- Can we measure how much prior errors bias subsequent token predictions?
- How does temporal anchoring maintain the learning signal in self-rewarding loops?
- Does higher temperature sampling help bootstrapping loops generate more training data?
- Why does reasoning catalyst data remain stable across multiple self-improvement iterations?
- What makes deliberate practice on your own errors more effective than copying others?
- What makes a self-supervised pruning metric work without labels at scale?
- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- How does STaR's correctness filtering extend to tasks without ground truth answers?
- How does Goodhart's Law apply to proxy rewards in self-training systems?