Line of inquiry
Inquiring lines›How should we train models for cap…›What systematic failures and vulne…›this line of inquiry
What are the consequences of models training on synthetic data?
A broader line of inquiry — a family of 35 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 35
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can models learn to generate their own training examples effectively?
- Does self-generated training data reduce a model's capability diversity?
- Why does self-generated training data outperform externally curated domain examples?
- What happens when models train on AI-generated content recursively?
- Why does self-generated training data outperform externally sourced data?
- How does the ratio of synthetic to real training data affect model collapse?
- How does diversity loss in synthetic data mirror tail distribution disappearance?
- What failure modes emerge when model-generated content trains on itself iteratively?
- Why do unified models still inherit data-distribution biases from training?
- Can trained models encode programs more complex than their data-generating process?
- Can a world model have rich representations without adequate data coverage?
- Can synthetic data generation work without seed examples?
- What causes irreversible model collapse when training on model-generated content?
- Can synthetic data preserve the diversity needed for transcendence to work?
- Can world models form from aggregated partial information across training distributions?
- What makes policy self-distillation more effective than external teacher distillation?
- Why does separating global coverage from local variation improve synthetic data generation?
- Can models detect statistical properties of their own generation in real time?
- What makes seed data a bottleneck in synthetic generation pipelines?
- What training data contamination rates threaten model safety most practically?
- What distinguishes instance seeds from full input-output exemplar requirements?
- Can the serving loop itself become the primary training data source?
- Does model collapse occur across different architectures or only in specific conditions?
- How does self-distillation differ from standard fine-tuning approaches?
- Why does the same training data produce different gains across models?
- How does training data distribution determine what models can learn?
- Can deterministic computation actually create new information in data?
- Why does reasoning catalyst data remain stable across multiple self-improvement iterations?
- Can seedless generation maintain explainability while scaling control?
- Can population-level distributions shift usefully even when individual prediction fails?
- How can smaller models help select useful data for larger models?
- Can synthetic data generation balance all three QDC axes simultaneously?
- Does debiasing training data actually solve the bias problem in machine learning?
- What output distribution properties make smaller models better for wide sampling?
- How does off-policy data reuse inside trust regions affect convergence guarantees?