Line of inquiry
Inquiring lines›How can we optimize language model…›How do reasoning capabilities emer…›this line of inquiry
What causes language models to reason excessively or insufficiently relative to problem difficulty?
A broader line of inquiry — a family of 28 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 28
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can models overthink and underthink at the same time?
- Why do language models overthink simple questions when given extra time?
- Why do models overthink easy problems and underthink difficult ones?
- Why do models overthink underspecified problems instead of rejecting them?
- Does distillation from reasoning models spread overthinking to smaller models?
- Can penalizing reasoning transitions fix underthinking without fine-tuning models?
- Do reasoning models overthink ill-posed questions instead of recognizing incompleteness?
- When does extended thinking hurt performance on easier problems?
- What triggers overthinking versus underthinking in reasoning models?
- Why does overthinking degrade performance at extreme recursion depths?
- Can reasoning models reject ill-posed questions or do they overthink?
- What happens when models overthink during test-time search?
- Why do different model training approaches produce different overthinking thresholds?
- Can extended deliberation in agents become counterproductive like human overthinking?
- Why do simple math problems get worse with longer reasoning chains?
- Why does inference-time thinking hurt proactive critical thinking in vanilla models?
- Can runtime confidence signals detect when reasoning has crossed the overthinking threshold?
- Do search agents face their own overthinking threshold like reasoning models do?
- How can we turn reasoning model failures into useful training signals?
- Why does extended thinking increase output variance without improving reasoning quality?
- Can conditioning generation on difficulty probes reduce overthinking on simple tasks?
- Can bounded workspaces prevent overthinking better than summarization alone?
- Do iterative refinement methods reproduce the same overthinking failure mode?
- How does compressing memory between iterations prevent overthinking?
- How does flip-event regression differ from premature thought path abandonment?
- How does self-distillation degrade reasoning by suppressing uncertainty signals?
- Is premature decision-making a form of underthinking in transformer models?
- Can a single model implement fast thinking, slow thinking, and tool use?