Theme of inquiry
When and why does chain-of-thought reasoning succeed or fail?
A question within its area, explored through 4 lines of inquiry below — each a family of specific questions the research asks.
43 specific questions
- Do chain-of-thought explanations reveal genuine reasoning or trigger latent features?
- Does chain-of-thought reasoning amplify bullshit or just make it more visible?
- Does CoT reasoning actually cause the outputs that follow it?
- Can chain-of-thought explanations be both sufficient and necessary for model decisions?
- Why does chain-of-thought fail when problems lack matching training schemata?
- Why do chain-of-thought prompts work if reasoning is not systematic?
- What three factors actually drive chain of thought performance improvements?
23 specific questions
- Why do corrupted reasoning traces sometimes generalize better than correct ones?
- Why do corrupted traces maintain performance as well as correct traces?
- Why do deliberately corrupted reasoning traces sometimes generalize better than correct ones?
- Do corrupted reasoning traces teach something different than pure success traces?
- Can corrupted reasoning traces be reliably distinguished from correct ones?
- What makes some reasoning traces better supervision than others despite equal accuracy?
- Why do invalid prompts produce reasoning traces as effectively as valid ones?
64 specific questions
- Can reasoning traces prove models are actually reasoning versus mimicking?
- Why do reasoning traces fail to accurately reflect model decision-making?
- Can reasoning traces serve purposes beyond producing the final answer itself?
- What makes a reasoning trace causally sufficient versus merely stylistically plausible?
- Why do models rarely admit to their actual reasoning in chain-of-thought traces?
- Why do reasoning traces resemble mimicry rather than verified problem-solving?
- Why do we measure reasoning quality by reading visible chains?
47 specific questions
- Why do correct reasoning traces tend to be shorter than incorrect ones?
- Do correct reasoning traces tend to be shorter than incorrect ones?
- Why do correct reasoning traces in language models tend to be shorter?
- Why do correct reasoning traces appear shorter than incorrect ones?
- Do longer chain-of-thought traces improve interpretability or just performance?
- Why are incorrect reasoning traces longer than correct ones?
- Why are shorter reasoning traces more reliable than longer correct ones?