Why doesn't showing an AI worked examples of 'fixing a mistake' teach it to catch its own mistakes later?
Why do SFT datasets fail to teach models error correction from correction traces?
This explores why fine-tuning a model on examples of mistakes being fixed doesn't produce a model that can actually catch and fix its own mistakes, and what the corpus says works better.
This explores why showing a model worked examples of errors being corrected, through supervised fine-tuning (SFT), doesn't reliably give it the ability to correct its own errors. The most direct answer in the corpus is a mismatch problem: the mistakes in a pre-collected correction dataset aren't the mistakes the model will actually make at test time. Models trained this way also tend to collapse into a single habitual 'correction move' instead of learning to diagnose what went wrong. What does work is online reinforcement learning across multiple turns, where the model practices fixing its own real errors as they happen Why does self-correction training on offline data fail?.
This is one example of a broader pattern: SFT is very good at teaching what a good answer looks like and much weaker at teaching the process that produces one. On optimization problems, SFT made outputs well formatted (valid JSON, correct sections) without making the solutions actually workable Does supervised fine-tuning actually improve reasoning on optimization problems?. On domain tasks, it raised final-answer accuracy while cutting how informative the reasoning was by nearly 39%, as models leaned on pattern-matching shortcuts Does supervised fine-tuning actually improve reasoning quality?. Seen this way, a correction trace is just another surface pattern to copy. The model learns that answers sometimes contain a 'wait, let me fix that' step, not how to notice that something is actually wrong.
The most surprising evidence for this is that reasoning traces may not need to be correct at all. Models trained on deliberately corrupted, irrelevant traces solved problems about as well as models trained on correct ones, which suggests traces work more like scaffolding than lessons Do reasoning traces need to be semantically correct?. If the content of a trace matters that little, a model learning from correction traces may be picking up their shape without their logic. The reverse also holds: some traces that are *right* still hurt fine-tuning, for example when the reasoning keeps exploring after the answer is already settled Does every correct chain-of-thought trace improve fine-tuning?. With SFT, a trace being correct doesn't mean it teaches well.
There is also a reason why practicing on your own errors matters so much. When a model's earlier mistakes are sitting in its context, its later error rate rises sharply, and making the model bigger doesn't fix this Do models fail worse when their own errors fill the context?. Recovering from your own contaminated context is a different skill from reading a clean story about someone else's recovery. The approaches that work all place the model inside its own failures. One RL method keeps varied failed trajectories as negative examples while filtering successful ones for quality, so the model learns what not to tolerate Why do correct code trajectories teach models to tolerate errors?. Even without any training, getting a model to make mistakes on its few-shot examples and then write down the principles it learned improves performance Does learning from mistakes improve in-context learning?.
The fix isn't just 'switch from SFT to RL', though. When SFT on expert data is followed by RL, training tends to go through a disruption, a readaptation and then overfitting, and blending the two objectives dynamically works better than running them one after the other Why does SFT-then-RL training follow a predictable three-phase pattern?. RL brings its own trap too: on problems that are far too hard, rare lucky successes get heavily rewarded and teach shortcuts instead of real reasoning Do overly hard RLVR samples actually harm model capabilities?. The takeaway is that error correction is learned by practicing on mistakes the model actually makes, at a difficulty where real fixes are possible, and not by copying examples of fixes.
Sources 10 notes
SFT on offline correction traces fails because training errors don't match test errors and models collapse into single correction modes. Multi-turn online RL under the model's own error distribution successfully trains self-correction by letting models practice correcting their actual mistakes.
Supervised fine-tuning makes model outputs look correct—proper JSON structure, valid identifiers, expected sections—without making them physically feasible. The model learns surface features of solutions, not the reasoning to construct valid ones.
SFT improves final-answer accuracy but reduces reasoning informativeness by 38.9% on average. Models reach correct answers through pattern-matching shortcuts rather than genuine inferential reasoning, becoming less auditable despite higher accuracy scores.
Models trained on systematically irrelevant traces maintain solution accuracy and sometimes improve out-of-distribution generalization, suggesting traces function as computational scaffolding rather than meaningful reasoning steps.
Post-conclusion reasoning—where the model keeps exploring after sufficient evidence for the answer—degrades supervised fine-tuning despite preserving correctness. Removing only this tail improves learning more than removing equally-long random suffixes, proving the harm comes from unnecessary exploration, not length.
Show all 10 sources
Error accumulation in context causes non-linear performance degradation in long-horizon tasks. Model scaling does not fix this; only test-time compute through thinking models reduces the effect by preventing error-contaminated context from biasing reasoning.
GRPO-RoC filters positive trajectories for quality while preserving diverse failures as negative signal, allowing a 14B model to reach frontier math performance in 510 RL steps, surpassing much larger models with cleaner reasoning.
LEAP demonstrates that models achieve better performance on reasoning and math tasks by intentionally erring on few-shot examples, reflecting on mistakes, and deriving explicit task-specific principles—without additional labeled data or fine-tuning.
CHORD identifies three distinct training phases: initial capability disruption from policy shift, readaptation to expert patterns, then overfitting. Dynamically weighting SFT as an auxiliary objective within on-policy RL resolves this progression and improves stability.
Training on nearly-impossible problems causes models to learn degenerate shortcuts rather than genuine reasoning, and these shortcuts contaminate pre-existing capabilities. Group-relative normalization treats rare accidental successes as high-advantage trajectories, reinforcing answer repetition and computation-skipping instead of sound reasoning patterns.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Large Language Models Cannot Self-Correct Reasoning Yet
- Learning to Reason for Factuality