Theme of inquiry
How do reasoning capabilities emerge and evolve across model scales?
A question within its area, explored through 10 lines of inquiry below — each a family of specific questions the research asks.
39 specific questions
- Why do correct reasoning traces tend to be shorter than incorrect ones?
- Do correct reasoning traces tend to be shorter than incorrect ones?
- Why are incorrect reasoning traces longer than correct ones?
- Why do correct reasoning traces in language models tend to be shorter?
- Why do correct reasoning traces appear shorter than incorrect ones?
- Do longer chain-of-thought traces improve interpretability or just performance?
- Why are shorter reasoning traces more reliable than longer correct ones?
31 specific questions
- Does chain of thought reasoning faithfully reflect what a model actually believes?
- Does chain-of-thought text causally drive reasoning or merely reflect it?
- Why do models rarely admit to their actual reasoning in chain-of-thought traces?
- Can chain-of-thought traces be faithful without causal sufficiency and necessity?
- Do chain-of-thought explanations reveal genuine reasoning or trigger latent features?
- Can chain-of-thought faithfulness exist without causal necessity in reasoning?
- Why do we measure reasoning quality by reading visible chains?
32 specific questions
- How does semantic association differ from mechanistic causal reasoning?
- Why do causal reasoning directions succeed while temporal reasoning directions fail?
- Do LLMs show stronger reasoning about causality than about temporal ordering?
- Can LLMs reason through semantics without understanding causal mechanisms?
- Can external actions provide causal necessity that language models lack?
- How might human-LLM teams reinforce each other's causal reasoning mistakes?
- What are collider structures and why do they reveal reasoning errors?
31 specific questions
- What three factors actually drive chain of thought performance improvements?
- What makes diffusion chain-of-thought reasoning qualitatively different from sequential chain-of-thought?
- Does chain-of-thought reasoning amplify bullshit or just make it more visible?
- What happens to chain-of-thought performance across distribution shifts?
- Why does chain-of-thought fail when problems lack matching training schemata?
- Does chain-of-thought reasoning specifically improve performance on metalinguistic tasks?
- How much of chain-of-thought reasoning is actually redundant?
37 specific questions
- Why do invalid reasoning steps produce nearly the same performance gains?
- Why do invalid prompts produce reasoning traces as effectively as valid ones?
- Why do verbalized reasoning chains fail on certain problem classes?
- Why do logically invalid chain-of-thought examples work nearly as well?
- Why does chain-of-thought prompting fail to fix length-induced reasoning degradation?
- Why do chain-of-thought prompts work if reasoning is not systematic?
- Can chain-of-thought traces harm rather than help user understanding?
26 specific questions
- Can reflection in reasoning models be corrective rather than just confirmatory?
- Why does reflection in reasoning models stay confirmatory instead of corrective?
- Why does reflection in reasoning models confirm rather than correct initial directions?
- Why does reflection in reasoning models mostly confirm the first answer?
- Does thought consolidation address the confirmatory reflection problem in reasoning models?
- Does reflection actually correct errors or just rationalize existing outputs?
- Why does reflection in reasoning models tend to be confirmatory rather than corrective?
47 specific questions
- Does thinking-token overuse actually degrade reasoning accuracy in practice?
- How does reasoning accuracy degrade when token budgets exceed critical thresholds?
- How do thinking tokens exhibit diminishing returns beyond a critical threshold?
- Can thinking token density explain reasoning performance beyond total length?
- Do tokens beyond a critical threshold actually improve reasoning quality?
- What happens to reasoning accuracy when models use more thinking tokens?
- What happens to model reasoning accuracy as thinking token requirements exceed critical thresholds?
40 specific questions
- Why does representation recycling of MI-peak tokens improve reasoning accuracy?
- Can layer-wise prediction stabilization identify when genuine reasoning has stopped?
- What sparse mechanistic structures drive reasoning traces in language models?
- How does reward density during training affect token efficiency in reasoning?
- What distinguishes genuine reasoning activation from memorization-assisted answer recall?
- Why does the first generated token trigger collapse of task superposition?
- Do reflection tokens and symbolic tokens serve different roles in reasoning?
28 specific questions
- Can models overthink and underthink at the same time?
- Why do language models overthink simple questions when given extra time?
- Why do models overthink easy problems and underthink difficult ones?
- Why do models overthink underspecified problems instead of rejecting them?
- Does distillation from reasoning models spread overthinking to smaller models?
- Can penalizing reasoning transitions fix underthinking without fine-tuning models?
- Do reasoning models overthink ill-posed questions instead of recognizing incompleteness?
21 specific questions
- Why does additional reasoning effort not improve theory of mind performance?
- Why do reasoning models perform poorly at theory of mind tasks?
- Why does reasoning volume fail to improve theory of mind performance?
- Why does increasing reasoning not improve AI social reasoning performance?
- Why does reasoning effort fail to improve theory of mind performance?
- Do longer reasoning traces actually improve theory of mind accuracy?
- Why do reasoning models perform worse on theory of mind tasks?