INQUIRING LINE

AI models often learn more from seeing new kinds of examples than from seeing more of the same kind — why?

Why does diversity of training cases matter more than raw dataset size?

This explores why a model learns more from a varied set of training examples than from simply piling on more examples, and what the collection says about how variety gets lost.


This explores why a varied set of training examples can teach a model more than a larger pile of similar ones. The short answer from the collection: once a model has seen enough examples of one kind, more of the same teaches it almost nothing. What it still lacks is coverage of situations it hasn't seen. The clearest evidence comes from data pruning. If you rank training examples by how hard they are and throw away the easy, repetitive ones, half of a standard image dataset can be removed with no loss in accuracy Can we prune training data without hurting model performance?. Pruning also changes the usual pattern where each doubling of data buys a smaller gain. Picking examples carefully can make performance improve much faster than adding data at random. So the dataset's raw size was never the useful part. Much of it was the same lesson repeated.

It also helps to know what diversity does that quality doesn't. One study of synthetic training data separates three properties that are usually blended together. Quality helps a model on problems like the ones it trained on. Diversity helps it handle problems unlike anything it trained on. Complexity strengthens both How do quality, diversity, and complexity affect synthetic data differently?. Most evaluation collapses these into one 'quality' score, so a pipeline can look like it's improving while it steadily loses the variety that lets the model generalize. When a model generates its own training data in a loop, that loss can't be undone. A related finding gives another reason to want variety: a model trained on many imperfect experts can outperform every one of them. Their mistakes are uncorrelated, so training averages them out, much like a majority vote Can models trained on many imperfect experts outperform everyone?. That only works if the experts really differ. A thousand copies of one expert's habits give you nothing to cancel out.

This is why the 'Artificial Hivemind' result matters. When researchers compared more than 70 language models on open-ended prompts, the models gave strikingly similar answers, largely because they share training data and alignment methods Do different AI models actually produce diverse outputs?. The internet is huge, but if everyone trains on the same huge thing, scale doesn't produce diversity. Surprisingly, if you want a model to generate varied synthetic data, smaller models around 500M parameters produce more distinct outputs per sample than larger ones. Bigger models pile their probability onto a few favorite answers Why aren't bigger models better for generating diverse outputs?.

The same pressure appears during training, not just in the dataset. Reinforcement learning tends to lock onto one output format from pretraining within the first pass and suppress the others Does RL training collapse format diversity in pretrained models?. Rewarding only correct final answers narrows the model's behavior even on problems it hasn't solved yet Does outcome-based RL diversity loss spread across unsolved problems?. Training can push back. Rewarding semantic diversity directly made outputs more varied and also better, on both creative and math tasks Can diversity optimization improve quality during language model training?. Step-by-step critique during self-training keeps a model from settling on one solution too early Do critique models improve diversity during training itself?.

The takeaway you might not expect is that diversity isn't a fixed property of a dataset you collect once. Models erode it as they train, and the processes that look like improvement (optimizing a reward, filtering for quality, training on your own outputs) are often what erode it. Collecting more data doesn't fix this; protecting variety does. One caveat: the collection mostly covers this through synthetic data, data pruning and reinforcement learning. It has less on pretraining datasets compiled by people.


Sources 9 notes

Can we prune training data without hurting model performance?

Research shows that ranking training examples by difficulty (EL2N, forgetting, memorization) and removing easy ones beats power-law scaling laws. On CIFAR-10, 50% of data was pruned without accuracy loss, and self-supervised metrics scaled the approach to ImageNet.

How do quality, diversity, and complexity affect synthetic data differently?

Quality drives in-distribution generalization, diversity enables out-of-distribution generalization, and complexity strengthens both. Current evaluation methods collapse these into a single quality metric, causing self-improvement loops to degrade through irreversible diversity loss.

Can models trained on many imperfect experts outperform everyone?

Models trained on diverse experts converge on consensus behavior that outperforms individuals. Low-temperature sampling concentrates outputs on this majority-voted consensus, denoising uncorrelated biases and errors across the training set.

Do different AI models actually produce diverse outputs?

INFINITY-CHAT analyzed 70+ models across 26K open-ended queries and found an "Artificial Hivemind" effect: models independently generate strikingly similar or identical responses due to overlapping training data and alignment procedures, undermining the diversity benefits of model ensembles.

Why aren't bigger models better for generating diverse outputs?

Research shows that for synthetic data generation, models around 500M parameters outperform larger ones in output diversity per sample. Larger models concentrate probability mass on preferred outputs, reducing the variety of distinct samples generated within a fixed budget.

Show all 9 sources
Does RL training collapse format diversity in pretrained models?

Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.

Does outcome-based RL diversity loss spread across unsolved problems?

RL that rewards only final answer correctness sharpens the policy globally, concentrating probability mass on correct trajectories for solved problems while simultaneously reducing diversity on unsolved ones. Historical exploration (training diversity via UCB-style bonuses) and batch exploration (test-time diversity via repetition penalties) require structurally different mechanisms.

Can diversity optimization improve quality during language model training?

DARLING jointly optimizes for quality and semantic diversity using a learned classifier, finding that diversity rewards catalyze exploration and produce higher-quality outputs than quality-only baselines across both creative and mathematical tasks.

Do critique models improve diversity during training itself?

Step-level critique in the training loop counteracts tail narrowing and maintains solution diversity across self-training iterations. This training-time benefit—preventing premature convergence—is more fundamental than test-time accuracy gains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.