Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data

Paper · arXiv 2404.01413 · Published April 1, 2024
Frontier AI Risk & RSI

The proliferation of generative models, combined with pretraining on webscale data, raises a timely question: what happens when these models are trained on their own generated outputs? Recent investigations into modeldata feedback loops proposed that such loops would lead to a phenomenon termed model collapse, under which performance progressively degrades with each model-data feedback iteration until fitted models become useless. However, those studies largely assumed that new data replace old data over time, where an arguably more realistic assumption is that data accumulate over time. In this paper, we ask: what effect does accumulating data have on model collapse? We empirically study this question by pretraining sequences of language models on text corpora. We confirm that replacing the original real data by each generation’s synthetic data does indeed tend towards model collapse, then demonstrate that accumulating the successive generations of synthetic data alongside the original real data avoids model collapse; these results hold across a range of model sizes, architectures, and hyperparameters. We obtain similar results for deep generative models on other types of real data: diffusion models for molecule conformation generation and variational autoencoders for image generation. To understand why accumulating data can avoid model collapse, we use an analytically tractable framework introduced by prior work in which a sequence of linear models are fit to the previous models’ outputs. Previous work used this framework to show that if data are replaced, the test error increases with the number of model-fitting iterations; we extend this argument to prove that if data instead accumulate, the test error has a finite upper bound independent of the number of iterations, meaning model collapse no longer occurs. Our work provides consistent empirical and theoretical evidence that data accumulation avoids model collapse.

Introduction. The advent of large-scale generative models such as GPT-4 (Achiam et al., 2023), DALL-E (Ramesh et al., 2022) and Stable Diffusion (Rombach et al., 2022) has revolutionized the field of artificial intelligence. These models, trained on vast web-scale datasets, exhibit remarkable capabilities in generating text, images, and other media (Brown et al., 2020; Saharia et al., 2022). However, as these models become more widely used, an increasing amount of generated data populates the web. This raises a critical question: what are the consequences of training generative models on datasets containing their own outputs?

Recent studies have investigated this question, revealing that training generative models on their own outputs can cause the performance of such models to progressively degrade with each model-fitting iteration, eventually rendering newer models useless (Hataya et al., 2023; Mart ́ınez et al., 2023a; Shumailov et al., 2023; Alemohammad et al., 2023; Mart ́ınez et al., 2023b; Bertrand et al., 2023; Briesch et al., 2023; Dohmatob et al., 2024a;b) (see Appendix A for review and discussion of prior work). This phenomenon was consequently labeled model collapse. Model collapse warns that democratizing access to generative models runs the risk of polluting the very data necessary to train future iterations of generative models.

To better understand this phenomenon many prior works have considered a setup that assumes each model’s generated data replaces previous data. In theory, this leads to very natural comparisons across generations as the total number of training points for each model remains fixed. In practice, subsequent generations of LLMs are often trained with increasing data over time – e.g., 1.4 trillion tokens for Llama 1 (Touvron et al., 2023a), 2 trillion for Llama 2 (Touvron et al., 2023b), 15 trillion for Llama 3 – in which presumably both human-generated and machine-generated data are accumulating in training sets collected from the internet. It was noted in some of those works Hataya et al. (2023); Mart ́ınez et al. (2023a); Alemohammad et al. (2023); Bertrand et al. (2023); Dohmatob et al. (2024b) that model collapse can be either slowed down or negated by mixing in clean data with the generated data.

To that end, in this work we study the effect of accumulating data on model collapse, rather than replacing data. Our data-accumulating setting is, in some sense, maximally pessimistic: it considers a hypothetical future where synthetic data are uncontrollably dumped on the internet to be vacuumed up for training the next iteration of generative models. Nevertheless, we find that model collapse is avoided when accumulating data.

We begin by studying model collapse experimentally with deep generative models trained on realistic data: transformers on causal language modeling (Sec. 2.1), diffusion models on molecular conformation (Sec. 2.2) and variational autoencoders on images (Sec. 2.3). After confirming that replacing data at every iteration indeed causes test error to increase with the number of iterations, we empirically find that accumulating synthetic data with real data avoids model collapse for all models and for all data modalities we test. To understand why replacing data and accumulating data have different consequences for model collapse, we turn to an analytically tractable framework of a sequence of linear models, each trained on synthetic outputs generated from the previous-iteration’s fitted linear model (Mobahi et al., 2020; Dohmatob et al., 2024a). Within this framework, Dohmatob et al. (2024a) demonstrated that if data are replaced with each model-fitting iteration, the test error increases linearly with the number of iterations n. We extend Dohmatob et al. (2024a)’s analysis to prove that if data instead accumulate, then the test error has a finite and (to us, surprisingly) well-controlled upper bound independent of the number of model-fitting iterations.1 Altogether, our work suggests that data accumulation may be robust to model collapse and emphasizes the importance of considering accumulating data and other real-world data dynamics in the analysis of model collapse in generative models trained on web-scale data.

Method. We first investigate model collapse experimentally in several classes of generative models. Here, and for the remainder of this manuscript, the term model collapse refers to notably worsening error over increasing iterations of the model-data loop, while avoiding model collapse refers instead to bounded error over such iterations. To test the effect of accumulating data on model collapse, we compare accumulating data against replacing data. We use three diverse experimental setups of causal transformers, diffusion models, and variational autoencoders trained on real text, molecular conformation, and image datasets, respectively. We find that replacing data yields model collapse for all models and all datasets, whereas accumulating data avoids model collapse.

Experiments We first train causal transformers (Vaswani et al., 2017) on text data. Specifically, we pretrain 9M parameter GPT-2 (Radford et al., 2019) and 12M, 42M and 125M parameter Llama2 (Touvron et al., 2023b) language models for a single epoch on TinyStories (Eldan & Li, 2023), a 470M token GPT-3.5/4-generated dataset of short stories at a kindergarten reading level. For each model-fitting iteration n ≥2, we sample a new dataset of the same size as TinyStories from the previous iteration’s language model and then either replace or concatenate the previous dataset with the newly generated dataset. In each model-fitting iteration, we then pretrain a newly initialized model on the replaced or concatenated dataset from the previous iteration. We experiment with sampling the new datasets using temperatures 0.3 or 1.0. We chose this combination of architectures, scales, dataset, and sampling because the setup necessitates pretraining multiple iterations Results We found that for all architectures, parameter counts, and sampling temperatures, as the number of model-fitting iterations increased, replacing data led to an increase in test cross entropy (Fig. 2 top). We also found that for all architectures, parameter counts, and sampling temperatures, as the number of model-fitting iterations increased, accumulating data led to equal-or-lower test cross entropy (Fig. 2 bottom). Lower temperature (0.3) led to a faster increase in test error than higher temperature (1.0) (Appendix Fig. 13), but the trend was consistent for both temperatures. Table 1 shows samples of generated texts for GPT2 (9M) and Llama2 (125M) models at model-fitting iterations 3-5 when both accumulating and replacing data, as well as iterations 8-10 (replacing only).

Ablations We ablate for several additional potential confounds beyond generation temperature. First, when accumulating data, subsequent model iterations are trained on larger datasets than when replacing data. To control for this, we also perform experiments in which data is replaced, but the size of the (fully synthetic) dataset is grown to match the training set size in the accumulation regime. We find that model performance still degrades (albeit at a lower rate). This is shown in Appendix C, Table 2, right-most column. Second, a possible concern could be that degrading performance when replacing data could be due to low model performance in iteration 1 (and thus the quality of the first synthetic dataset). We control for this by varying the amount of training performed in iteration 1 only and find that this has no significant impact. Lastly, we find that our results are also consistent across varying dataset sizes and training epochs. These ablations are discussed in Appendix F.

Experiments We next train sequences of diffusion models on molecular conformation data. Specifically, we train GeoDiff (Xu et al., 2022), a geometric diffusion model for molecular conformation generation, on the GEOM-Drugs (Axelrod & Gomez-Bombarelli, 2022) dataset. We down-sample the training split of GEOM-Drugs to 40, 000 molecular conformations, which we use as our initial training set, and perform 50 diffusion steps for each prediction. For the loss, we use the standard loss used by GeoDiff: a weighted variational lower bound to the conditional likelihood; for more details, see Xu et al. (2022).

Discussion. This work explored the phenomenon of model collapse, an important concern as AIgenerated content permeates the internet and finds its way into future training datasets. Prior work has shown that training on model outputs can lead to degraded performance (Mart ́ınez et al., 2023a;b; Shumailov et al., 2023; Alemohammad et al., 2023; Hataya et al., 2023; Bertrand et al., 2023; Briesch et al., 2023; Dohmatob et al., 2024a;b), implying that future model training faces a difficult challenge of ensuring strict training dataset hygiene. For a significantly more thorough discussion of related work, please see Appendix A.

Our findings extend these prior works to show that if data accumulates and models train on a mixture of “real” and synthetic data, model collapse no longer occurs. We show this both experimentally on causal transformers for language modeling, diffusion models for molecule generation, and variational auto-encoders on image data as well as theoretically for linear regression. Together, these results strongly suggest that the “curse of recursion” may not be as dire as had been portrayed – provided we accumulate synthetic data alongside real data, rather than replacing real data by synthetic data only.

Looking to the future, many questions worth investigating remain. For instance, in future work we would like to explore different data generation and accumulation regimes, such as (1) additional “real” data being introduced in each model-fitting iteration and (2) different schedules of how much synthetic data is generated at each iteration and (3) human-filtering of generated data, e.g., as done in RLHF. Additionally, we note that in all our experiments, the synthetic dataset is generated by sampling from the previous model, i.e., with some stochasticity; in future work, we would like to explore also what happens if data is generated deterministically, e.g. with temperature 0 in a typical language model.

Lastly, it is worth noting that “model collapse” – as a term of art – has been used in various ways by various researchers; so care is required in comparing claims across articles. In reviewing the literature, we identified at least four related phenomena: (0) unbounded test error blowup (as here); (1) modal collapse — collapse to one (or a few) modes; (2) collapse to uniformity; and (3) amplification of artifacts introduced by models fit to previous synthetic data. Future work should map out what factors cause which to occur and what preventative strategies are effective at addressing each.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do models learn from self-generated outputs without cascading failures? How do training data quality and composition affect downstream model performance? How do interpretive frames override surface features in text comprehension? How do hallucinated citations emerge in AI scholarly output? Does training data format shape model reasoning more than domain content? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? Why do training associations persist despite contradictory contextual information? How do curriculum design and feedback approaches affect model learning? How reliably can humans and AI detectors identify machine-generated text? How effectively can test-time voting aggregate diverse reasoning samples? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? How does diversity prevent model convergence on superficial patterns?