Theme of inquiry
How can training approaches develop reasoning beyond surface imitation?
A question within its area, explored through 7 lines of inquiry below — each a family of specific questions the research asks.
71 specific questions
- Does scaling reasoning capability create tradeoffs with instruction following?
- Does fine-tuning push models toward reasoning shortcuts that bypass the chain entirely?
- Do reasoning models switch approaches when encountering local difficulty?
- Does reasoning fine-tuning actually damage a model's ability to abstain?
- Does reasoning fine-tuning actually harm a model's ability to abstain?
- Can models reason at inference without specialized internal training?
- Does reasoning fine-tuning actually reduce a model's ability to abstain?
42 specific questions
- Does thinking-token overuse actually degrade reasoning accuracy in practice?
- How do thinking tokens exhibit diminishing returns beyond a critical threshold?
- How does reasoning accuracy degrade when token budgets exceed critical thresholds?
- What happens to model reasoning accuracy as thinking token requirements exceed critical thresholds?
- Does a critical thinking token threshold exist for model accuracy?
- Why does reasoning accuracy degrade beyond a critical thinking token threshold?
- Can thinking token density explain reasoning performance beyond total length?
75 specific questions
- Why might latent reasoning capture types of thinking that verbalized CoT cannot?
- Can latent reasoning achieve the same substitution without tokens?
- Is chain-of-thought reasoning actual computation or distribution imitation?
- Can latent reasoning scale test-time compute without verbalized tokens or special training?
- Can steering a single latent feature replicate chain-of-thought performance?
- Can reasoning happen in latent space without chain of thought?
- When does explicit reasoning actually degrade performance on a task?
19 specific questions
- Why does training data format shape reasoning strategy more than domain content?
- How much does training data format influence reasoning strategy versus domain content?
- How does training format shape reasoning strategy more than content?
- Why does training data format shape reasoning strategy more than content?
- Does training data format shape reasoning strategy more than domain content?
- Can training format itself shape what reasoning strategy a model learns?
- Does training data format shape model reasoning more than domain content?
75 specific questions
- Why do reasoning gains resist clear attribution to specific training changes?
- Why do instruction following and reasoning capability trade off in training?
- Can smaller amounts of diverse reasoning demonstrations replace exhaustive factual training data?
- Can small demonstration sets unlock general reasoning without large question data?
- Does task diversity in pretraining data transfer reasoning better than larger models?
- How do single training examples activate reasoning capabilities in language models?
- How much does pre-training frequency predict reasoning task performance?
17 specific questions
- Why does additional reasoning effort not improve theory of mind performance?
- Why do reasoning models perform poorly at theory of mind tasks?
- Why does increasing reasoning not improve AI social reasoning performance?
- Why does reasoning volume fail to improve theory of mind performance?
- Why does reasoning effort fail to improve theory of mind performance?
- Why might social reasoning work differently than formal logical reasoning?
- Do longer reasoning traces actually improve theory of mind accuracy?
56 specific questions
- Why do benchmark scores rise while reasoning quality declines?
- Why do AI benchmarks measure accuracy instead of reasoning quality?
- How can high benchmark performance mask broken reasoning in AI systems?
- Can benchmark improvements hide degradation of deliberative reasoning?
- Do reasoning benchmarks predict real performance in long delegated workflows?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- Should benchmarks measure trace length or whether constraints were actually satisfied?