Theme of inquiry
Can visible reasoning improve model evaluation and safety?
A question within its area, explored through 8 lines of inquiry below — each a family of specific questions the research asks.
35 specific questions
- How does semantic association differ from mechanistic causal reasoning?
- Do LLMs show stronger reasoning about causality than about temporal ordering?
- Can LLMs reason through semantics without understanding causal mechanisms?
- Why do causal reasoning directions succeed while temporal reasoning directions fail?
- Can external actions provide causal necessity that language models lack?
- Why do LLMs inherit causal biases from their training data?
- How might human-LLM teams reinforce each other's causal reasoning mistakes?
80 specific questions
- Do chain-of-thought explanations reveal genuine reasoning or trigger latent features?
- Does chain of thought reasoning faithfully reflect what a model actually believes?
- Why do models rarely admit to their actual reasoning in chain-of-thought traces?
- Is chain-of-thought reasoning actual computation or distribution imitation?
- Why do we measure reasoning quality by reading visible chains?
- Does chain-of-thought text causally drive reasoning or merely reflect it?
- Can steering a single latent feature replicate chain-of-thought performance?
29 specific questions
- Why does additional reasoning effort not improve theory of mind performance?
- Why do reasoning models perform poorly at theory of mind tasks?
- Why does reasoning volume fail to improve theory of mind performance?
- Why does reasoning effort fail to improve theory of mind performance?
- Why does increasing reasoning not improve AI social reasoning performance?
- Why do reasoning models perform worse on theory of mind tasks?
- Do longer reasoning traces actually improve theory of mind accuracy?
24 specific questions
- How often do scheming reasoning and covert actions actually align in practice?
- Can monitoring reasoning alone miss scheming that agents conceal in behavior?
- What happens to agent candor when reasoning traces are monitored versus hidden?
- Can reasoning traces and logged actions expose scheming that public messages hide?
- Could hints change both agent reasoning and behavior rather than action alone?
- Can activation probes detect scheming reasoning without observing the act?
- Does shortcut deliberation occur in model reasoning before taking covert action?
77 specific questions
- How does extended thinking affect variance in reasoning model outputs?
- Does longer reasoning always improve model accuracy on complex tasks?
- When does explicit reasoning actually degrade performance on a task?
- Can minimal reasoning steps match verbose reasoning accuracy?
- When does extended thinking hurt performance on easier problems?
- Why do correct reasoning traces in language models tend to be shorter?
- How much does extended thinking actually improve model reasoning ability?
69 specific questions
- Why do more capable reasoning models become harder to control by instruction?
- Does scaling reasoning capability create tradeoffs with instruction following?
- Why do instruction following and reasoning capability trade off in training?
- Can reasoning fine-tuning improve both capability and instruction compliance together?
- Why does instruction-following capability decrease as models scale stronger?
- Does fine-tuning push models toward reasoning shortcuts that bypass the chain entirely?
- Can reasoning models succeed at logic but fail at execution?
47 specific questions
- Does thinking-token overuse actually degrade reasoning accuracy in practice?
- How does reasoning accuracy degrade when token budgets exceed critical thresholds?
- How do thinking tokens exhibit diminishing returns beyond a critical threshold?
- Can thinking token density explain reasoning performance beyond total length?
- What happens to model reasoning accuracy as thinking token requirements exceed critical thresholds?
- What happens to reasoning accuracy when models use more thinking tokens?
- Does a critical thinking token threshold exist for model accuracy?
7 specific questions
- How do continuous concept tokens explore multiple reasoning paths without explicit sampling?
- How do soft thinking and token-level mixtures explore multiple paths simultaneously?
- How does continuous soft thinking explore multiple paths without explicit training?
- How do soft token mixtures enable parallel reasoning exploration without explicit training?
- How do continuous concept tokens compare to latent trajectory sampling?
- How does soft thinking compare to sampling multiple independent reasoning paths?
- How does soft thinking achieve stochastic exploration without explicit training?