INQUIRING LINE

Does adding more training data always help an AI model, or can too much of the wrong data make it worse?

What determines whether additional training data still improves model performance?

This explores when adding more training data stops helping a model, or starts hurting it, and what decides which side of that line you're on.


This explores when more training data stops paying off, and what decides whether the next batch helps, does nothing, or does harm. The corpus gives a clear answer: the amount of data matters less than how well the data fits the model learning from it. A pile of examples that is larger but badly matched can leave a model worse off than a small, well-chosen slice.

The most surprising evidence is that mixed datasets often contain examples that actively work against the skill you care about. In one study, picking just 5% of an instruction dataset, chosen because its examples push the model in the same direction as the target task, beat training on all of it. The other examples were not just extra weight. They pulled the model's reasoning style away from what the task needed Can we train better models on less data?. Image classification shows the same pattern from a different angle. Removing the easy, redundant examples let half of a dataset be dropped with no loss in accuracy, and returns improved much faster than the usual 'add more data, get slowly diminishing gains' curve predicts Can we prune training data without hurting model performance?.

Difficulty cuts both ways. Examples that are too easy teach nothing new, but examples that are too hard can be worse than useless. When models are trained with reinforcement learning on problems they almost never solve, the rare lucky successes get heavily rewarded. The model then learns shortcuts like repeating answers or skipping steps, and these habits damage skills it already had Do overly hard RLVR samples actually harm model capabilities?. One response is to let the model set its own difficulty level. In self-play setups, one copy of the model writes problems sized to what the other copy can almost solve, which produces a curriculum that keeps pace with the learner and needs no outside data Can language models improve themselves without any external training data?.

'Fit' also covers style as well as difficulty. Data that a stronger teacher model has polished can lower a student model's performance when the improvements go beyond what the student is ready to absorb. The student does better when it keeps only the refinements that match its own patterns Does teacher-refined data always improve student model performance?. Taking this further, models sometimes learn more from training data they rewrote themselves than from data a stronger model wrote for them. Self-generated data raised question-answering accuracy from 33.5% to 47%, likely because the model puts information into a form it can actually use Does self-generated training data improve model learning?.

Finally, there is a ceiling where the question changes. Once one model's pretraining has saturated, extra compute does more good spent on training a diverse group of models and combining their predictions than on feeding the same model more passes over the data Should extra compute refine one model or build many?. The takeaway you might not have expected: 'more data' is the wrong unit to think in. What matters is whether each example sits at the edge of what this particular model can learn. Once a single model can't move further, the next gain may come from adding more models, not more data.


Sources 7 notes

Can we train better models on less data?

LESS uses low-rank gradient features to select instruction data most similar to target capabilities, and training on the selected 5% consistently outperforms full dataset training. The improvement occurs because mixed datasets contain examples that actively hinder specific skills by shifting reasoning strategy away from task requirements.

Can we prune training data without hurting model performance?

Research shows that ranking training examples by difficulty (EL2N, forgetting, memorization) and removing easy ones beats power-law scaling laws. On CIFAR-10, 50% of data was pruned without accuracy loss, and self-supervised metrics scaled the approach to ImageNet.

Do overly hard RLVR samples actually harm model capabilities?

Training on nearly-impossible problems causes models to learn degenerate shortcuts rather than genuine reasoning, and these shortcuts contaminate pre-existing capabilities. Group-relative normalization treats rare accidental successes as high-advantage trajectories, reinforcing answer repetition and computation-skipping instead of sound reasoning patterns.

Can language models improve themselves without any external training data?

SQLM uses a proposer-solver framework where the proposer generates calibrated problems and the solver learns via majority-vote verification. Both agents improve through RL alone, creating an automatic curriculum that scales without human labels or ground-truth answers.

Does teacher-refined data always improve student model performance?

Teacher-refined data degrades performance when it exceeds the student's learning frontier, even if objectively higher quality. Students should filter refinements using their own statistical profile to retain only compatible improvements.

Show all 7 sources
Does self-generated training data improve model learning?

SEAL demonstrates that models learn better from synthetic data they generate themselves than from data created by stronger external models. Self-generated data improved QA performance from 33.5% to 47.0%, suggesting that model-specific restructuring aligns with the learner's representational needs.

Should extra compute refine one model or build many?

Once single-model pretraining saturates, aggregating predictions from a diverse population of models reaches lower validation loss than further refining one model. Anti-correlated learning-rate and weight-decay schedules plus chain distillation enable this efficiently, matching 256-epoch ensembles with ~56 epochs.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.