AI models often say 'wait, let me double-check' — but does that second-guessing ever actually change their answer?
Why does reflection in reasoning models rarely overturn initial answers?
This explores why the 'wait, let me double-check' passages in reasoning models like o1 and DeepSeek-R1 so seldom change the model's final answer, and what that says about what reflection is actually doing.
This explores why the self-checking passages in reasoning models (the 'wait, let me reconsider' moments) so rarely flip the answer, and what that reveals about what reflection really does. The short version from the corpus: in practice, reflection mostly confirms an answer the model has already settled on. An analysis of eight reasoning models found that reflections rarely change the initial answer and mainly work as after-the-fact confirmation Is reflection in reasoning models actually fixing mistakes?. The surprising part is what training does. When you train models on longer reflection chains, the gain shows up in how often the *first* answer is right, not in the model's ability to catch its own mistakes. Because the confirmation adds so little, you can stop generation early, saving about a quarter of the tokens for under 3% accuracy loss Does reflection in reasoning models actually correct errors?.
That doesn't make reflection tokens meaningless. Words like 'Wait' and 'Therefore' are moments where the model's internal state carries unusually strong information about the correct answer. Suppressing them hurts accuracy, and suppressing the same number of random tokens doesn't Do reflection tokens carry more information about correct answers?. Together, these findings point to an interesting split: reflection matters, but as a point where the model consolidates its reasoning rather than as an audit that overturns it. The answer is largely decided before the double-check begins.
Why doesn't real correction happen more often? Real correction needs specific skills: noticing that an assumption was wrong, backtracking, and rebuilding from that point. Benchmarks that test these skills directly expose the gap. On constraint-satisfaction puzzles that force genuine backtracking, top reasoning models score only about 20–23% Can reasoning models actually sustain long-chain reflection?. When reflection is broken down into assumption-making, backtracking and self-refinement, models trained on reasoning traces fall apart wherever revision is actually needed What makes reflection actually work in reasoning models?. A related line of work argues that reasoning traces are partly stylistic imitation: invalid steps perform almost as well as valid ones Do reasoning traces show how models actually think?. If models learned the *look* of self-checking from their training data, they can produce convincing 'wait…' passages without the machinery to act on what those passages say.
The failures run in both directions, too. Sometimes models drop promising paths too early. They hop between ideas mid-exploration, and simply penalizing these switches at decoding time improves accuracy without any retraining Do reasoning models switch between ideas too frequently? Why do reasoning models abandon promising solution paths?. Other times they cling to the framing they were handed. Models go along with false premises in a question even when they demonstrably know the correct facts Why do language models accept false assumptions they know are wrong?, and reasoning models will churn out long answers to unanswerable questions instead of saying the question is broken Why do reasoning models overthink ill-posed questions?. Neither pattern is disciplined revision: one gives up on good ideas, the other never questions the starting point.
The takeaway you might not have expected: more thinking isn't the same as more self-correction. Accuracy peaks at a middling chain-of-thought length, and more capable models prefer *shorter* chains Why does chain of thought accuracy eventually decline with length?. Some apparent reasoning collapses turn out to be failures to carry out long procedures in text, and these disappear once the model can use tools Are reasoning model collapses really failures of reasoning?. So if you want a model that truly changes its mind, the corpus suggests longer reflection won't get you there. You would need to train backtracking and premise-checking directly, or give the model outside checks (tools, verifiers) that can actually contradict it.
Sources 12 notes
Analysis of 8 reasoning models shows reflections rarely change answers and primarily serve as post-hoc confirmation. Training on longer reflection chains improves first-answer quality, not self-correction capability.
Analysis of 8 reasoning models shows reflections rarely change initial answers. Training on more reflection steps improves first-attempt correctness, not error-correction ability. Early stopping saves 24.5% tokens with only 2.9% accuracy loss.
Specific tokens like "Wait" and "Therefore" show sharp spikes in mutual information with correct answers. Suppressing them harms reasoning while suppressing equal random tokens does not, and representation recycling improves accuracy 20%.
DeepSeek-R1 and o1-preview achieve only 20-23.6% exact match on 850 constraint satisfaction problems requiring genuine backtracking. This ceiling reveals that reflective reasoning fluency does not translate to actual problem-solving competence on unfamiliar instance structures.
LR²Bench decomposes reflection into three measurable capabilities: assumptions, backtracking, and self-refinement. Models trained on reasoning traces collapse at tasks requiring actual constraint-satisfying revision, suggesting current reflection training improves surface fluency, not genuine correction.
Show all 12 sources
LLM reasoning traces perform as persuasive appearances rather than reliable explanations of computation. Invalid logical steps perform nearly as well as valid ones, and corrupted traces generalize comparably, showing that semantic correctness is not what produces the performance gains.
o1-like models frequently abandon reasoning paths mid-exploration, wasting tokens on incomplete approaches. A decoding-only penalty on thought-transition tokens (TIP strategy) discourages switching, improving accuracy on challenging math without model fine-tuning.
Reasoning LLMs exhibit two reinforcing failures: wandering (invalid exploration) and underthinking (premature path-switching). Decoding-level interventions like thought-switching penalties improve accuracy without fine-tuning, suggesting viable solutions exist but are abandoned prematurely.
The FLEX Benchmark shows that models reject false presuppositions at rates far below acceptable levels (GPT-4: 84%, Mistral: 2.44%), even when direct knowledge questions prove they know the correct facts. False presuppositions drive more accommodation than correct knowledge drives rejection.
Reasoning models generate redundant, lengthy responses to questions with missing premises while non-reasoning models correctly identify them as unanswerable. Training optimizes for producing reasoning steps but never teaches models when to disengage.
Task accuracy peaks at intermediate CoT length, with optimal length increasing alongside task difficulty but decreasing with model capability. RL training naturally gravitates toward shorter chains as models improve, revealing that simplicity emerges from reward signals rather than explicit training.
Models confined to text-only generation cannot execute multi-step procedures at scale, even when they know the underlying algorithm. Tool-enabled models solve problems beyond the supposed reasoning cliff, suggesting the bottleneck is procedural execution bandwidth.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- First Try Matters: Revisiting the Role of Reflection in Reasoning Models
- Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
- When Hindsight is Not 20/20: Testing Limits on Reflective Thinking in Large Language Models
- Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
- Large Language Models Cannot Self-Correct Reasoning Yet
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap