Line of inquiry
Inquiring lines›Where does language-model reasonin…›How do reward models guide reliabl…›this line of inquiry
Why does self-revision increase model confidence while degrading accuracy?
A broader line of inquiry — a family of 30 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 30
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does self-revision in reasoning chains amplify confidence in wrong answers?
- Why does model self-revision increase confidence while degrading accuracy?
- How does self-revision on wrong answers increase model confidence further?
- Does internal self-revision actually degrade reasoning accuracy in models?
- Why do reasoning models amplify confidence in incorrect answers during self-revision?
- Why does single-model self-revision amplify confidence in incorrect answers?
- Why do reasoning models struggle with self-evaluation and revision?
- Why does single-agent self-revision amplify confidence in wrong answers over time?
- Do self-revision tokens measurably degrade reasoning accuracy in scaled models?
- How do self-revisions degrade reasoning accuracy in extended traces?
- Why does self-revision degrade reasoning accuracy in o1-like models?
- Does self-revision actually improve reasoning in large language models?
- Does external critique guide revision better than internal self-assessment during model training?
- Why do reasoning models exhibit self-doubt about their own early assessments?
- Can single models correct their own beliefs without amplifying confidence in wrong answers?
- Why does most refinement in iterative models maintain answers rather than improve them?
- Why does self-critiquing actually reduce plan quality in language models?
- Do reasoning models need to verbalize doubt to correct their own mistakes?
- Why does revision often make reasoning accuracy worse in frontier models?
- Why does external critique improve revision while internal self-assessment fails?
- Why does iterative refinement amplify rather than correct reasoning errors?
- Why does external critique improve revision accuracy more than self-assessment?
- Can external retrieval signals outperform internal self-assessment during revision?
- Can debate between multiple models prevent the failures of single-model self-revision?
- Why does uncontrolled self-revision drift toward instance-specific overfitting?
- Can a model evaluate its own improvements without degrading over iterations?
- Why do final answers contradict what the thinking draft explicitly concluded?
- Why does systematic overconfidence on self-generated outputs compound autoregressive errors?
- How does self-distillation degrade reasoning by suppressing uncertainty signals?
- Why do models confabulate inconsistently across different samples?