Line of inquiry
Inquiring lines›How do training choices shape mode…›How do training dynamics and archi…›this line of inquiry
How does training on self-generated data affect model capabilities?
A broader line of inquiry — a family of 56 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 56
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can models learn to generate their own training examples effectively?
- Why does self-generated training data outperform externally curated domain examples?
- Does self-generated training data reduce a model's capability diversity?
- Why does self-generated training data outperform externally sourced data?
- What happens when models train on AI-generated content recursively?
- Why do unified models still inherit data-distribution biases from training?
- Can trained models encode programs more complex than their data-generating process?
- Why do proprietary models improve with training while open-source models decline?
- Do correlated human errors prevent models from transcending their training sources?
- How does self-distillation differ from standard fine-tuning approaches?
- What makes policy self-distillation more effective than external teacher distillation?
- What failure modes emerge when model-generated content trains on itself iteratively?
- When does knowledge distillation produce student models superior to teachers?
- Can world models form from aggregated partial information across training distributions?
- How do learning dynamics on one example shift predictions on other responses?
- How does the Learning Law explain why all examples should contribute equally?
- Why do weaker models generate better training data than stronger models?
- Does the model learn depth-wise drift as an explicit strategy?
- How much do structural inductive biases matter compared to training data volume?
- Why does the same training data produce different gains across models?
- Can the serving loop itself become the primary training data source?
- How do instruction backtranslation and MAGPIE demonstrate self-generation principles?
- How does distribution mismatch between training and deployment break self-correction?
- Can unsupervised confidence-based training scale to domains beyond human evaluation reach?
- How does training data distribution determine what models can learn?
- Can we predict out-of-distribution generalization without access to downstream tasks?
- Do newer language model generations improve forecasting ability without additional training?
- Can synthetic data generation work without seed examples?
- Which AI imaginaries dominate training data and shape system behavior most strongly?
- Can we cheaply estimate which samples are currently most informative?
- What distinguishes instance seeds from full input-output exemplar requirements?
- Why do models fail under distribution shift if accuracy metrics stay high?
- How do unstated constraints become invisible to training data distributions?
- Can diverse expert demonstrations exceed the knowledge of any single expert?
- Can population-level distributions shift usefully even when individual prediction fails?
- Why does reasoning catalyst data remain stable across multiple self-improvement iterations?
- Why is latent-level prediction more sample-efficient than token-level prediction?
- How does correctness emergence occur when no expert initially solved the task?
- Can deterministic computation actually create new information in data?
- Why does monological training prevent models from overriding statistical priors?
- How do labs actually train next-generation models from previous ones?
- Can ensemble predictions be distilled back into a single deployable model?
- Why does filtering for correct examples prevent error compounding in self-training?
- Why do optimal learning dynamics improve scaling law coefficients specifically?
- Can self-distillation reduce catastrophic forgetting in continual learning?
- Does unpredictable generalization from SDF become predictable at different training document scales?
- What makes a self-supervised pruning metric work without labels at scale?
- How does joint backpropagation differ from training separate ensemble models?
- Why do energy-based models generalize better on out-of-distribution data than standard transformers?
- Why are post-cutoff test sets essential for evaluating genuine forecasting ability?
- Can seedless generation maintain explainability while scaling control?
- What makes data augmentation an implicit form of contraction learning?
- How do RL subnetworks identified from different random seeds compare?
- Can predictive self-supervision work on unlabeled sequential visual data?
- Why does pure numeric ID indexing force models to learn from scratch?
- How do training data cutoffs produce false claims that stay consistent?