Line of inquiry
Inquiring lines›How can we optimize language model…›How do test-time resources and tra…›this line of inquiry
Do reasoning and instruction-following trade off as models scale?
A broader line of inquiry — a family of 64 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 64
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does scaling reasoning capability create tradeoffs with instruction following?
- Why do models learn reasoning form instead of actual abstract inference?
- Why do instruction following and reasoning capability trade off in training?
- Does reasoning structure match explicit versus implicit task demands?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- Why does instruction-following capability decrease as models scale stronger?
- Can models learn when to think versus answer directly?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- Can models learn to select exemplars based on reasoning skills rather than complexity?
- Can models learn when to invoke search during reasoning tasks?
- Do models genuinely reason harder on difficult tasks or just appear to?
- Do higher asymptote recipes unlock genuinely novel reasoning strategies?
- Does adding reasoning to models degrade other capabilities like rule inference?
- Can models be trained to explain instead of imitate answers?
- Why does explicit theory injection work better than example-based learning for reasoning tasks?
- Why do foundation models develop heuristics instead of world models?
- Can explicit constraint statements override the dominance of surface heuristics?
- Can format adaptation alone explain why reasoning enrichment improves instruction following?
- How does scaling reasoning capability actually reduce instruction-following ability?
- Why do models fail on logically equivalent tasks with different data distributions?
- Can surface heuristics override implicit constraints in domain-specific reasoning?
- Can tools unlock reasoning strategies that require abstract insight beyond computation?
- Why does distillation transfer reasoning patterns with few examples?
- Can scaffolding frameworks isolate inductive reasoning from deductive confounds?
- Do task-specific heuristics improve gradually or appear suddenly at scale?
- Does iterative computation for reasoning transfer to environment dynamics modeling?
- How do foundation models develop task-specific heuristics instead of world models?
- What distinguishes task-specific heuristics from genuine world models?
- What changes when reasoning models adopt trajectory-response output formats?
- What inductive bias would force models to learn Newtonian mechanics instead of shortcuts?
- What distinguishes coherent reasoning from inaccurate but plausible predictions?
- Can models learn to optimize their own chain-of-thought generation?
- Does the heuristic dominance ratio vary predictably across model architectures?
- Why does imitation learning create a ceiling for reasoning capability?
- What makes a causal abstraction more transferable than a generic heuristic?
- Do reasoning models trade instruction following for deliberative capability?
- Why does stronger reasoning reduce model compliance with instructions?
- How do search and reasoning workflows improve forecasting performance over base models?
- Why do human-curated thought examples fail to improve model thinking?
- Can instruction tuning succeed without explicit task understanding?
- How do humans and R1 models differ in information gain patterns?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- What distinguishes inductive inference from negative evidence versus positive patterns?
- What distinguishes surface mechanisms from the training regimes that produce them?
- Why do familiar patterns that support correct answers sometimes drive errors?
- Can models maintain multiple task interpretations simultaneously before committing to a single policy?
- What distinguishes conceptual understanding from statistical pattern matching in models?
- Does reasoning style transfer matter more than solution correctness in distillation?
- How does inductive reasoning from partial evidence enable hypothesis formation?
- What metric distinguishes deep reasoning from superficial information propagation?
- What makes training-free approaches like Soft Thinking preferable to SoftCoT?
- Can machine learning encode pragmatic reasoning about when rules should bend?
- What real-world forecasting domains benefit most from contextual reasoning integration?
- Does next-state prediction alone build mechanistic world models or just sophisticated interpolation?
- How does contrapositive augmentation change the tractability of reasoning tasks?
- Why do foundation models develop task-specific heuristics instead of causal understanding?
- How do unstated feasibility constraints affect model decision-making?
- How does o1-style reasoning relate to learned search processes versus memorized solutions?
- What test-time strategies did o3 discover without human specification?
- Why do macro and micro forecasting scales require different reasoning approaches?
- What happens to safety guardrails when we scale reasoning without instruction control?
- How does treating cognition as computation reshape education and work?
- Which constraint types do reasoning models handle best?
- Can reasoning style be steered as a single linear direction?