SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

Does model collapse depend on how we schedule training data?

Does replacing real data with synthetic data each generation cause inevitable model collapse, or is collapse avoidable through different training schedules? This matters because it determines whether training on generated content is fundamentally limited.

Synthesis note · 2026-10-08 · sourced from Frontier AI Risk & RSI

Gerstgrasser et al. test whether model collapse is an artifact of how prior studies built their training loops rather than an inherent property of training on synthetic data. Prior work assumed "each model's generated data replaces previous data" — a setup the authors call realistic-sounding but actually inconsistent with practice, since real LLM training sets grow across generations (1.4 trillion tokens for Llama 1, 2 trillion for Llama 2, 15 trillion for Llama 3) rather than swapping old data for new. Pretraining sequences of causal transformers (GPT-2 9M; Llama2 12M/42M/125M) on TinyStories, and resampling synthetic data from each generation's model, they report: "we confirm that replacing the original real data by each generation's synthetic data does indeed tend towards model collapse," while "accumulating the successive generations of synthetic data alongside the original real data avoids model collapse." The same contrast holds for GeoDiff diffusion models on molecular conformations and VAEs on images, "across a range of model sizes, architectures, and hyperparameters."

The mechanism comes from an analytically tractable linear-regression framework (extending Dohmatob et al., 2024a) in which a sequence of linear models is fit to the previous model's outputs. Under replacement, prior theory showed test error grows linearly with the number of fitting iterations. The authors "extend this argument to prove that if data instead accumulate, the test error has a finite upper bound independent of the number of iterations" — collapse is a property of the replacement schedule, not an inevitable consequence of training on model-generated text. Ablations narrow the claim further: growing a synthetic-only dataset to match the accumulated dataset's size still degrades performance, just more slowly, and varying how much training happens in the first iteration has no significant effect. So the result isn't simply "more data helps" — it's that keeping the original real data present, generation after generation, is what bounds the error.

This sits in direct tension with Does training on AI-generated content permanently degrade model quality?, which frames collapse as irreversible tail-loss and already flags — without detailing — "Gerstgrasser et al. 2024" as counter-evidence that the debate isn't settled; this note is that paper. The two findings aren't describing the same experiment: the tail-loss note's irreversibility holds under a replace-data regime (the Shumailov et al. setup); this paper shows that irreversibility is conditional on that regime and disappears once generated data is concatenated with, rather than substituted for, the real corpus. It also complicates How do quality, diversity, and complexity affect synthetic data differently?: where QDC locates generalization effects in properties of the synthetic data itself (diversity, complexity), this paper locates the collapse/no-collapse outcome in the accumulation schedule — whether real data stays in the mix — rather than in any property of the synthetic data generated.

The study is small-scale and deliberately synthetic: kindergarten-level TinyStories text, models up to 125M parameters, and a "maximally pessimistic" assumption that synthetic data is dumped uncontrollably with no filtering or curation. It does not test real web-scale pretraining mixtures, does not address deterministic (temperature-0) generation (left to future work), and only resolves one of four phenomena the authors themselves group under "model collapse" — unbounded test-error blowup — leaving modal collapse, collapse to uniformity, and artifact amplification untouched. The implication the evidence supports is conditional: the curse of recursion is avoidable in principle as long as real human-generated data keeps accumulating alongside synthetic output rather than being crowded out by it — a bound that depends on real data continuing to be collected, not merely having existed historically.

Inquiring lines that read this note 4

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do models learn from self-generated outputs without cascading failures? How do training data quality and composition affect downstream model performance?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 110 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

model collapse is avoided when synthetic data accumulates alongside real data instead of replacing it