Give an AI model the same amount of training data, but make it AI-written instead of real — why does it actually get worse?
Why does synthetic-only data degrade performance even when matched in size?
This explores why training a model only on AI-generated data tends to make it worse than training on the same amount of real data, and what the synthetic data is missing beyond sheer volume.
This explores why AI-generated training data falls short of real data even when you give the model just as many examples. One caveat first: the corpus has no head-to-head experiment that holds size fixed and swaps real data for synthetic. What it does have is a set of explanations for what synthetic data tends to lack. Together they suggest that the size of a dataset was never the right thing to count.
The clearest answer is that "quality" is really three properties that do different jobs. How do quality, diversity, and complexity affect synthetic data differently? separates them. Quality helps a model on problems like the ones it trained on. Diversity helps it handle problems it hasn't seen. Complexity strengthens both. Most evaluation squeezes these into a single score, so a synthetic dataset can score well and still be narrow. When models train on their own outputs over and over, the variety drains away and doesn't come back. A synthetic set the same size as a real one can cover far less ground. Real data spans many more situations, and that range is what synthetic generation tends to lose.
The second answer is about realism, meaning whether examples fit together the way real ones do. Why does random tool sampling produce unrealistic synthetic training data? shows this in data for teaching models to use tools. Randomly paired tools would never be used together in practice, and single question-and-answer examples ignore how real conversations unfold over several turns. Every example can look fine on its own while the set as a whole teaches patterns that don't exist. Fixing this took structure (sampling tools that are actually related, planning whole dialogues), not more data.
The third answer is that synthetic data has limits on what it can teach at all. Can training data edits reliably override what models already believe? finds that synthetic documents reliably add new facts but are unreliable at changing what a model already believes. Does synthetic document finetuning fail at larger scales? shows the same problem in a safety setting: the effect is unpredictable, not just weak. A related pattern turns up in RL. Does RL training collapse format diversity in pretrained models? finds that training amplifies one dominant style and squeezes out the others. Generated data seems to pull the same way, toward the model's most typical outputs.
The surprising part is that the fix isn't simply "use real data." What makes synthetic data work across different domains and models? argues there is no universal recipe, because what matters changes with the domain and the model. The more promising work makes variety something you control on purpose. Can we generate synthetic data without any seed examples? first builds a map (a taxonomy) of everything the data should cover, then varies examples within each area. Can synthetic data replace seed examples in task generation? starts generation from small task pieces instead of copying full examples. When synthetic-only data falls short, the corpus points to missing variety and broken realism, not to having too few examples.
Sources 8 notes
Quality drives in-distribution generalization, diversity enables out-of-distribution generalization, and complexity strengthens both. Current evaluation methods collapse these into a single quality metric, causing self-improvement loops to degrade through irreversible diversity loss.
Random tool sampling fails because unrelated tools cannot credibly compose, and Q&A framing ignores multi-turn dialogue coherence. ToolFlow shows that sampling tools from relevance graphs and generating with dialogue plans closes this gap.
Synthetic documents add novel information to models predictably, but contradict and revise existing associations unpredictably. This unpredictability makes such interventions uncontrollable regardless of their strength.
Research shows SDF cannot reliably inoculate models against misalignment from reward hacking, with the effect unpredictable rather than merely weak. The conclusion is explicitly scoped to the specific scales tested, leaving extrapolation uncertain.
Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.
Show all 8 sources
Research shows no single optimal recipe for synthetic data generation. The impact of data properties like complexity and diversity varies by domain, model, use case, and scale, making explainable, flexible control more valuable than one-size-fits-all methods.
Simula separates global coverage from local diversity, using taxonomy construction for coverage and agentic refinement for complexity. This architecture makes all three desiderata—quality, diversity, complexity—controllable simultaneously without requiring seed data.
TarGEN generates synthetic data using atomic task elements (instance seeds) instead of full input-output examples, achieving 1-3 point improvements on SuperGLUE tasks. The approach works by constraining label generation after seeding inputs, enabling data creation for domains with no prior examples.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Reasoning-Driven Synthetic Data Generation and Evaluation
- Orchestrating Synthetic Data with Reasoning
- A Little Human Data Goes A Long Way
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
- ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
- CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks