SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

Can language models fix their own reasoning mistakes?

Do LLMs actually improve their answers when asked to reconsider them without external feedback? This question matters because many papers claim self-correction works, but the evidence may be misleading.

Synthesis note · 2026-10-08 · sourced from Frontier AI Risk & RSI

The paper defines "intrinsic self-correction" as an LLM revising its own answer "based solely on its inherent capabilities, without the crutch of external feedback," and tests it on reasoning with GPT-3.5-Turbo, GPT-4, GPT-4-Turbo, and Llama-2-70b-chat. The finding: "LLMs struggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correction." Prior papers claiming gains (Kim et al. 2023; Shinn et al. 2023) turn out to depend on oracle labels telling the model whether its answer was already right — "the improvements vanish when oracle labels are not available." A second artifact is prompt design: some claimed self-correction gains "stem from the sub-optimal prompt for generating initial responses," where the feedback step merely supplies instructions the first prompt should have included; folding those instructions into the initial prompt erases the advantage.

The paper's mechanism, in its own terms: when the model is already well-aligned and the initial prompt well-designed, "the initial response should already be optimal relative to the prompt and the specific decoding algorithm." A feedback prompt is then an additional, unhelpful input that can "bias the model away from producing an optimal response to the initial prompt." The paper also re-examines multi-agent debate (Du et al. 2023), where multiple LLM instances critique each other's answers, and replicates it on GSM8K at matched inference cost against self-consistency. Debate's gains are "no better than self-consistency... when considering an equivalent number of responses" — the paper argues debate is really "a means to achieve 'consistency' across multiple model generations," differing from self-consistency only in whether the vote is model-driven or count-based, so "the observed improvement is evidently not attributed to 'self-correction,' but rather to 'self-consistency.'"

This sharpens Is reflection in reasoning models actually fixing mistakes? and Does reflection in reasoning models actually correct errors?: those studies found reflection in trained reasoning models rarely overturns the initial answer; this paper gets the same directional result via a different mechanism — explicit multi-round self-correction prompting on non-reasoning-tuned chat models — and adds the methodological diagnosis (oracle-label leakage, unfair baselines) explaining why earlier work saw gains where none existed. It also parallels Does self-revision actually improve reasoning in language models? in showing degradation rather than improvement from revision, and gives a prompting-level counterpart to What limits how much models can improve themselves?: without an external verifier (oracle label, tool, or trained critic), the model has no advantage over its own generation to exploit, so correction has nothing to work with. Why does self-correction training on offline data fail? addresses the same gap at training time rather than at inference time.

The excerpt tests reasoning tasks only, on 2023-era chat models prompted explicitly to self-correct — it does not establish that self-correction fails for other domains; the authors note style, safety, and preference alignment are reported elsewhere as cases where LLMs "properly evaluate whether a response is inappropriate," unlike judging their own reasoning errors. It also does not test whether models trained specifically to produce useful reflection (rather than prompted post hoc) behave differently, leaving open whether the limitation is about prompting LLMs to self-correct or about self-correction as such.

Inquiring lines that read this note 5

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does self-revision amplify confidence in wrong model answers? How do models learn from self-generated outputs without cascading failures? How can evaluations be made robust against model reward hacking? How do users confuse explanation quality with actual system accuracy? What prevents LLMs from applying their reasoning knowledge to improve outputs?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 135 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llms cannot self-correct reasoning without oracle feedback — multi-agent debate works only as disguised self-consistency