Line of inquiry
Inquiring lines›How do training and design choices…›Can visible reasoning improve mode…›this line of inquiry
Why do stronger reasoning capabilities create tradeoffs with instruction following?
A broader line of inquiry — a family of 69 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 69
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do more capable reasoning models become harder to control by instruction?
- Does scaling reasoning capability create tradeoffs with instruction following?
- Why do instruction following and reasoning capability trade off in training?
- Can reasoning fine-tuning improve both capability and instruction compliance together?
- Why does instruction-following capability decrease as models scale stronger?
- Does fine-tuning push models toward reasoning shortcuts that bypass the chain entirely?
- Can reasoning models succeed at logic but fail at execution?
- Why does instruction tuning hurt knowledge-intensive tasks more than reasoning tasks?
- Why does latent reasoning override no-think instructions in models?
- How does scaling reasoning capability actually reduce instruction-following ability?
- Can models learn when to think versus answer directly?
- Why do strong models struggle more with instruction following than mid-tier ones?
- Does reasoning structure match explicit versus implicit task demands?
- Do models genuinely reason harder on difficult tasks or just appear to?
- Can instruction-level interventions fix memory-induced reasoning failures in practice?
- How does question difficulty and breadth affect what models learn to reason?
- Why do models learn reasoning form instead of actual abstract inference?
- Can models be trained to explain instead of imitate answers?
- Can models learn when to invoke search during reasoning tasks?
- Are instruction-tuned models more or less sensitive to prompt semantics than others?
- Why do difficult problems force models to develop reasoning strategies?
- Do reasoning models fail to report processes that actually influence their answers?
- Why does stronger reasoning reduce model compliance with instructions?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- How can we turn reasoning model failures into useful training signals?
- Why do reasoning models fail to improve constrained optimization performance?
- Do models trained for safety over-refuse compared to models trained for reasoning?
- Can we steer model reasoning by manipulating single features?
- Do higher asymptote recipes unlock genuinely novel reasoning strategies?
- Is the reasoning cliff actually a tool-use problem?
- Does decoupling reasoning from tool use actually improve accuracy?
- Can explicit constraint statements override the dominance of surface heuristics?
- Can training models on backward reasoning improve their forward planning ability?
- Why do models fail on logically equivalent tasks with different data distributions?
- How does learnability at the observer's current state prevent novelty from breaking model reasoning?
- Do reasoning models trade instruction following for deliberative capability?
- Why do reasoning-optimized models show no resistance advantage on agreement tasks?
- Can format adaptation alone explain why reasoning enrichment improves instruction following?
- Why does instruction-tuning reduce a model's context-following behavior?
- How do reasoning-related features behave when trained on near-impossible problems?
- Why do reasoning model failures stem from execution rather than reasoning?
- Can a model predict the right action but execute the wrong one?
- What changes when reasoning models adopt trajectory-response output formats?
- Why does reapplying the same computation stages improve model performance?
- Why do reasoning-optimized models show no sycophancy resistance advantage?
- Do shorter reasoning chains maintain instruction adherence better than longer ones?
- Can models distinguish between activated knowledge and genuine reasoning?
- Can instruction tuning succeed without explicit task understanding?
- Why must procedural skills consolidate before strategic reasoning can develop?
- Why does target probability matter more than task logical complexity?
- Why do human-curated thought examples fail to improve model thinking?
- Can models distinguish between logical impossibility and their own execution limits?
- Why does fine-tuning sometimes damage chain-of-thought reasoning even when accuracy improves?
- Can models maintain multiple task interpretations simultaneously before committing to a single policy?
- Why do familiar patterns that support correct answers sometimes drive errors?
- Does the heuristic dominance ratio vary predictably across model architectures?
- Can a single SAE feature control reasoning behavior across model families?
- How do unstated feasibility constraints affect model decision-making?
- Why does SFT fail when expert demonstrations are too long for small models?
- Can machine learning encode pragmatic reasoning about when rules should bend?
- Can models be trained to verify feasibility before proposing plans?
- Why do only two of fourteen models improve when problem constraints are removed?
- Why do models commit to answers early on easy versus hard tasks?
- Why do students learn better from explanations than from solving problems from scratch?
- Can a situationally aware model recognize and refuse planted shortcuts on purpose?
- How does belief-behavior inconsistency relate to instruction execution splits?
- How does the knowing-doing gap widen as tasks become more complex?
- Why do foundation models develop task-specific heuristics instead of causal understanding?
- Can you steer reasoning by directly manipulating SAE features?