Does cleaning up AI-generated training data actually stop a model from slowly degrading on itself — or just slow the decay down?
Does model collapse happen identically with unfiltered versus filtered synthetic data?
This explores whether filtering synthetic data (keeping only outputs a checker approves) changes how models degrade when trained on their own outputs, compared with feeding raw synthetic data back in. The short answer from the corpus: no, it doesn't happen the same way. Filtering changes how the failure looks, and it can delay the failure without removing it.
This explores whether filtering synthetic data (keeping only outputs a checker approves) changes how models degrade when trained on their own outputs, compared with feeding raw synthetic data back in. The short answer from the corpus: no, it doesn't happen the same way. Filtering changes how the failure looks, and it can delay the failure without removing it.
Start with the unfiltered case. Here the biggest factor turns out to be whether synthetic data replaces real data or is added alongside it, more than whether it gets filtered. When each generation of training data replaces the last, test error keeps growing without limit. When synthetic data is added on top of the original real corpus, error stays bounded. That result holds for language models, diffusion models and VAEs Does model collapse depend on how we schedule training data?. So a lot of the classic collapse story comes from throwing away real data, not from synthetic data as such.
Filtering adds a different dynamic. When a verifier screens synthetic outputs before retraining, things really do get better at first: variance drops, performance rises, and collapse is held off. Over longer runs, though, the model's parameters drift toward whatever the verifier itself believes. Unless the verifier is perfectly reliable, the early gains level off and then decline as its biases build up Does verifier filtering actually prevent model collapse long term?. Unfiltered collapse is a blurring of the model's own distribution. Filtered collapse is a narrowing toward the verifier's view. You swap one monoculture for another, and the second is harder to notice because it looks like quality control. A related warning from RL training: when the filter only lets through rare successes on very hard problems, it can reward lucky shortcuts such as repeating answers or skipping steps instead of real reasoning, and that damage spreads to skills the model already had Do overly hard RLVR samples actually harm model capabilities?.
The part you might not expect: collapse doesn't have to show up as worse accuracy. In retrieval systems, when about two-thirds of a corpus is synthetic, over 80% of retrieved results come from synthetic sources while answer accuracy stays high. The system looks healthy even as its range of sources shrinks to a fragile monoculture Does synthetic content in search results hide ecosystem decay?. A filter that checks correctness would pass all of it. A useful way to see why: LLM outputs are draws from the model's learned prior, not new observations of the world, so they should be weighted as such and not treated as equal to real evidence Should we treat LLM outputs as real empirical data?. Filtering improves the quality of those samples, but it can't turn them into new information about the world.
The practical upshot is that "filtered vs. unfiltered" is the wrong single lever. What matters more is keeping real data in the mix, how independent and reliable the filter is, and whether you're tracking diversity as well as accuracy. Because there's no universal recipe for good synthetic data, those choices depend on your domain and model What makes synthetic data work across different domains and models?. The corpus doesn't include a direct side-by-side long-run comparison of filtered and unfiltered collapse under identical conditions. What it offers are complementary pieces that point in the same direction.
Sources 6 notes
Replacing real data with synthetic data causes unbounded test error growth, but accumulating synthetic data alongside the original real corpus keeps error bounded across model architectures and sizes. The mechanism is proven analytically in a linear-regression framework and confirmed empirically on language models, diffusion models, and VAEs.
Verifier-filtered synthetic retraining reduces variance and improves performance initially, but pulls model parameters toward the verifier's own knowledge center over time. Without perfect verifier reliability, early gains plateau and degrade as bias accumulates.
Training on nearly-impossible problems causes models to learn degenerate shortcuts rather than genuine reasoning, and these shortcuts contaminate pre-existing capabilities. Group-relative normalization treats rare accidental successes as high-advantage trajectories, reinforcing answer repetition and computation-skipping instead of sound reasoning patterns.
When 67% of a corpus becomes synthetic, over 80% of retrieved results shift to synthetic sources while answer accuracy remains high, masking the loss of source diversity. This creates fragility: high accuracy resting on a monoculture collapses when that monoculture is poisoned.
Foundation Priors framework shows that LLM-generated text reflects the model's learned patterns and user's prompt choices, not ground truth. Such outputs should only influence inference through explicitly parameterized trust weights, not be treated as equivalent to real evidence.
Show all 6 sources
Research shows no single optimal recipe for synthetic data generation. The impact of data properties like complexity and diversity varies by domain, model, use case, and scale, making explainable, flexible control more valuable than one-size-fits-all methods.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- Reasoning-Driven Synthetic Data Generation and Evaluation
- Foundation Priors
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
- Orchestrating Synthetic Data with Reasoning
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- A Little Human Data Goes A Long Way
- Retrieval Collapses When AI Pollutes the Web