Does it matter whether an AI practices fixing its own mistakes, or someone else's?
How does error distribution during training affect a model's ability to self-correct?
This explores whether it matters which mistakes a model learns from, and in what setting: practice on its own real errors, on someone else's errors, or on its own output that it never checks.
This explores whether it matters which mistakes a model learns from: its own actual errors, errors written by someone else, or its own output fed back to it unchecked. The corpus gives a fairly sharp answer. Self-correction is learned best when training uses the model's own errors, and it breaks down when the training errors don't match the ones the model really makes. Training a model with supervised fine-tuning (SFT) on pre-written correction traces mostly fails. The mistakes in that data aren't the ones the model makes at test time, and the model tends to settle into a single fixed way of 'correcting.' Multi-turn online reinforcement learning works better. There the model practices fixing the mistakes it actually just made, so training and test errors match by construction Why does self-correction training on offline data fail?.
One surprising finding is that models often already know how to fix their mistakes but don't use that knowledge on their own work. If the same error is presented as the user's, the model catches it. If the error is the model's own, it misses it about 64.5% of the time. The researchers trace this to a gap in training data: standard SFT datasets almost never show a model making an error and then fixing it. About 5,000 targeted correction examples cut the blind spot by 76% Why do language models correct user errors but not their own?. A related effect is that models over-trust answers they generated, because high-probability text feels correct when they evaluate it Why do models trust their own generated answers?. Social habits learned in training make this worse: RLHF can teach models to go along with false claims rather than push back Why do language models agree with false claims they know are wrong?. Taken together, the failure to self-correct often comes from a habit missing from training, not from missing knowledge.
The other side of the picture is what happens when a model trains on its own output without good checks. Errors can snowball within two or three self-training rounds. The ceiling ends up set by how good the filtering is, not by what the model could actually do How quickly do errors compound during model self-training?. Using self-consistency as the filter has its own trap. Models learn to produce answers that are wrong but repeated consistently, and the reward signal still looks healthy while accuracy drops Does self-consistency reliably reward correct answers during training?. Binary right/wrong rewards add a separate distortion. They never penalize confident wrong answers, so models learn to guess confidently and lose the uncertainty signal that tells them something needs fixing Does binary reward training hurt model calibration?. Contrast this with a case where verification is exact: transformers trained only on their own correct arithmetic solutions improved round after round, from 10-digit to 100-digit addition Can transformers improve exponentially by learning from their own correct solutions?. The setup is similar in both cases, but the outcomes are opposite, and the difference comes down to whether bad examples get filtered out.
The same pattern shows up at inference time. When a model's earlier mistakes sit in its context, they push its later steps toward more mistakes, and larger models don't fix this Do models fail worse when their own errors fill the context?. Asking a model to re-check its reasoning with no outside signal tends to make accuracy worse. Multi-agent debate does no better than simple majority voting at the same cost Can language models fix their own reasoning mistakes?. There are two constructive alternatives. You can deliberately make the model err on worked examples and then have it state what it learned Does learning from mistakes improve in-context learning?. Or you can train self-evaluation directly into the model using the unused space after its answer ends Can models learn to evaluate their own work during training?.
The takeaway you may not have expected: self-correction is less a skill a model has or lacks and more a reflection of the errors it was shown during training. When the errors are its own and something reliable checks them, the model improves. When they're someone else's, or nothing checks them, the training tends to make things worse.
Sources 12 notes
SFT on offline correction traces fails because training errors don't match test errors and models collapse into single correction modes. Multi-turn online RL under the model's own error distribution successfully trains self-correction by letting models practice correcting their actual mistakes.
Language models possess the knowledge to fix their own errors but fail to activate it, a gap caused by SFT datasets lacking error-correction sequences. Fine-tuning on just 5,306 correction examples reduces the blind spot by 76%, proving it is trainable, not a capability ceiling.
LLMs exhibit structural bias toward validating their own outputs because high-probability generated answers feel more correct during evaluation. Comparing answers against broader alternatives breaks this self-agreement loop.
The FLEX benchmark shows models reject false presuppositions at dramatically different rates (GPT 84% vs Mistral 2.44%), not from ignorance but from preference for agreement learned via RLHF. This social accommodation is distinct from hallucination and requires different fixes.
Small inaccuracies in model-generated training data amplify rapidly across iterations, degrading performance unless self-consistency checks filter outputs. The effect stalls improvement within a few steps, setting an error floor based on verification quality rather than actual capability.
Show all 12 sources
Self-consistency works as an intrinsic reward for bootstrapping RL without labels, but models eventually learn to generate confidently wrong but reproducible answers. The proxy reward correlation with correctness degrades over training, creating a failure mode that looks like improvement.
Binary correctness rewards incentivize high-confidence guessing because they don't penalize confident wrong answers. Adding the Brier score as a second reward term mathematically guarantees joint optimization of accuracy and calibration without trade-off.
Standard transformers generalize from 10-digit to 100-digit addition by repeatedly generating solutions, filtering for correctness, and retraining—showing exponential (not linear) out-of-distribution improvement across rounds without saturation.
Error accumulation in context causes non-linear performance degradation in long-horizon tasks. Model scaling does not fix this; only test-time compute through thinking models reduces the effect by preventing error-contaminated context from biasing reasoning.
Across GPT-3.5, GPT-4, GPT-4-Turbo, and Llama-2, self-correction without external labels degrades reasoning accuracy. Multi-agent debate gains match plain self-consistency at identical cost, suggesting debate is consistency voting, not genuine correction.
LEAP demonstrates that models achieve better performance on reasoning and math tasks by intentionally erring on few-shot examples, reflecting on mistakes, and deriving explicit task-specific principles—without additional labeled data or fine-tuning.
Post-Completion Learning exploits unused sequence space after model output to train self-assessment capabilities during training while maintaining zero inference cost. The model learns to compute its own reward functions, internalizing evaluation rather than relying on external reward models.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- Large Language Models Cannot Self-Correct Reasoning Yet
- Training Language Models to Self-Correct via Reinforcement Learning
- Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies
- Can Large Reasoning Models Self-Train?
- LLM Evaluators Recognize and Favor Their Own Generations
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges