Does AI training on AI-made content break models because the data is bad, or because it crowds out real human data?
Is model collapse a property of data replacement or synthetic data itself?
This explores whether models degrade when trained on AI-generated data because the synthetic data is inherently harmful, or because of how it gets used, specifically when it pushes the original human-written data out of the training mix.
This explores whether model collapse comes from synthetic data being harmful in itself, or from the training practice of letting synthetic data push real data out. The clearest answer in the corpus points to replacement. When each new model generation is trained only on the previous generation's outputs, test error keeps growing without limit. When synthetic data is added on top of the original real data and the real data is kept, error stays bounded. This holds in a mathematical proof and in experiments on language models, image diffusion models and VAEs Does model collapse depend on how we schedule training data?. So the disaster scenario depends on the schedule. Synthetic data alone does not doom a model.
That doesn't make synthetic data neutral, though. One useful way to think about it: an LLM's output is not a fresh observation of the world. It is a sample from the model's own learned beliefs, shaped by whoever wrote the prompt Should we treat LLM outputs as real empirical data?. Training on it partly means training on your own assumptions. That's why replacement is so damaging. Once real data leaves the loop, nothing outside the model is left to correct its drift.
A common fix is to filter synthetic data through a verifier and keep only the outputs it approves. This helps at first, because quality goes up and noise goes down. Over time, though, the model gets pulled toward whatever the verifier believes, including its blind spots, and the early gains level off and then decline Does verifier filtering actually prevent model collapse long term?. Filtering doesn't remove the problem. It swaps the generator's biases for the verifier's.
The less obvious finding is that collapse-like narrowing doesn't need synthetic data at all. Chat models often give bland, repetitive answers, a problem called mode collapse. One line of research traces it to human preference data: annotators systematically favor familiar-sounding text, and training bakes that bias in Where does mode collapse in language models really come from?. In reinforcement learning, models can also fall back on generic, one-size-fits-all reasoning templates when the training signal stops telling good answers apart from bad ones Why do language models collapse into generic templates?. In both cases the narrowing comes from a feedback loop that stops supplying new, distinguishing information. The source of the data doesn't matter.
So the corpus suggests that 'synthetic versus real' is the wrong question. A better one is whether anything independent of the model still enters the training loop. Keeping real data, building diversity deliberately, and treating generated text as a weighted prior rather than as evidence all address that. How well synthetic data works also varies by domain, model and scale, so there's no universal safe recipe What makes synthetic data work across different domains and models?.
Sources 6 notes
Replacing real data with synthetic data causes unbounded test error growth, but accumulating synthetic data alongside the original real corpus keeps error bounded across model architectures and sizes. The mechanism is proven analytically in a linear-regression framework and confirmed empirically on language models, diffusion models, and VAEs.
Foundation Priors framework shows that LLM-generated text reflects the model's learned patterns and user's prompt choices, not ground truth. Such outputs should only influence inference through explicitly parameterized trust weights, not be treated as equivalent to real evidence.
Verifier-filtered synthetic retraining reduces variance and improves performance initially, but pulls model parameters toward the verifier's own knowledge center over time. Without perfect verifier reliability, early gains plateau and degrade as bias accumulates.
The research traces mode collapse to annotators' systematic preference for familiar text, a cognitive bias baked into training data. Verbalized Sampling, a training-free prompting method, restores 1.6–2.1× diversity by asking models to articulate probability distributions over responses.
When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.
Show all 6 sources
Research shows no single optimal recipe for synthetic data generation. The impact of data properties like complexity and diversity varies by domain, model, use case, and scale, making explainable, flexible control more valuable than one-size-fits-all methods.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
- Foundation Priors
- Reasoning-Driven Synthetic Data Generation and Evaluation
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
- When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs
- The Future of Facts: Tracing the Factual Generation-Verification Gap