Line of inquiry
Inquiring lines›How should agents manage and coord…›How can training approaches develo…›this line of inquiry
What capability tradeoffs emerge when scaling model reasoning abilities?
A broader line of inquiry — a family of 71 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 71
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does scaling reasoning capability create tradeoffs with instruction following?
- Does fine-tuning push models toward reasoning shortcuts that bypass the chain entirely?
- Do reasoning models switch approaches when encountering local difficulty?
- Does reasoning fine-tuning actually damage a model's ability to abstain?
- Does reasoning fine-tuning actually harm a model's ability to abstain?
- Can models reason at inference without specialized internal training?
- Does reasoning fine-tuning actually reduce a model's ability to abstain?
- Can reasoning fine-tuning improve both capability and instruction compliance together?
- Why does reasoning fine-tuning reduce a model's ability to abstain?
- Why does latent reasoning override no-think instructions in models?
- Why does instruction-following capability decrease as models scale stronger?
- Can models learn when to think versus answer directly?
- Why does step-by-step reasoning degrade performance on judgment-based tasks?
- Why does instruction tuning hurt knowledge-intensive tasks more than reasoning tasks?
- Does fine-tuning models for specific tasks destroy their ability to reason?
- Can models learn when to invoke search during reasoning tasks?
- Does adding reasoning to models degrade other capabilities like rule inference?
- How does scaling reasoning capability actually reduce instruction-following ability?
- Does reasoning structure match explicit versus implicit task demands?
- Can reasoning models succeed at logic but fail at execution?
- Why does reasoning fine-tuning reduce models' ability to abstain?
- Can penalizing reasoning transitions fix underthinking without fine-tuning models?
- Do models genuinely reason harder on difficult tasks or just appear to?
- Why do non-reasoning models work better under extreme decomposition than reasoning models?
- Are reasoning models more vulnerable to persuasion than standard models?
- Why does per-step deliberation lose global perspective compared to dynamic discovery?
- Can activation steering compress reasoning without retraining models?
- Why does inference-time thinking hurt proactive critical thinking in vanilla models?
- Can we detect redundant reasoning steps during model inference instead of training?
- Do higher asymptote recipes unlock genuinely novel reasoning strategies?
- Can training models on backward reasoning improve their forward planning ability?
- Does penalizing thought transitions improve reasoning without model retraining?
- Why do reasoning models fail to improve constrained optimization performance?
- How much reasoning depth do we actually need for most real-world tasks?
- Why does reasoning fine-tuning reduce model abstention capacity by 24 percent?
- Why does stronger reasoning reduce model compliance with instructions?
- Why does reasoning backward enable better forward reasoning performance?
- Can activation-space steering vectors replicate thinking model performance without retraining?
- Do models trained for safety over-refuse compared to models trained for reasoning?
- Can inflection points in reasoning detect when models genuinely change their minds?
- Why do strong models struggle more with instruction following than mid-tier ones?
- Is the reasoning cliff actually a tool-use problem?
- Why does strategy diversity within reasoning chains improve model generalization?
- Why do reasoning-optimized models show no resistance advantage on agreement tasks?
- Do reasoning models trade instruction following for deliberative capability?
- Can models learn to optimize their own chain-of-thought generation?
- Why does reapplying the same computation stages improve model performance?
- Does iterative computation for reasoning transfer to environment dynamics modeling?
- Can activation steering vectors compress reasoning without retraining models?
- Does this reasoning steering method work consistently across all model sizes?
- Why do foundation models develop heuristics instead of world models?
- What changes when reasoning models adopt trajectory-response output formats?
- What limits external scaling when a model lacks reasoning foundation?
- Why do reasoning models verbalize reasoning shortcuts less than necessary?
- How do reasoning-related features behave when trained on near-impossible problems?
- How should timing for reasoning intervention be determined during inference?
- How does chain-of-thought reasoning become decorative after domain-specific fine-tuning?
- Can you control LLM reasoning strategy without fine-tuning the model?
- How do models integrate conflicting signals in reasoning tasks?
- How do foundation models develop task-specific heuristics instead of world models?
- What makes training-free approaches like Soft Thinking preferable to SoftCoT?
- What distinguishes task-specific heuristics from genuine world models?
- Does the heuristic dominance ratio vary predictably across model architectures?
- Can activation steering directly steer models toward concise reasoning without prompting?
- Can single representation edits match chain-of-thought reasoning without explicit steps?
- Can a single model implement fast thinking, slow thinking, and tool use?
- What happens to safety guardrails when we scale reasoning without instruction control?
- Can you steer reasoning by directly manipulating SAE features?
- What test-time strategies did o3 discover without human specification?
- Why do foundation models develop task-specific heuristics instead of causal understanding?
- Can deterministic recurrent depth achieve the computational benefits of stochastic reasoning?