Line of inquiry
Inquiring lines›What determines reliable reasoning…›How does chain-of-thought reasonin…›this line of inquiry
Does scaling reasoning capability create fundamental tradeoffs in control and reliability?
A broader line of inquiry — a family of 85 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 85
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do reasoning models switch approaches when encountering local difficulty?
- Does scaling reasoning capability create tradeoffs with instruction following?
- Why do more capable reasoning models become harder to control by instruction?
- Can reasoning models succeed at logic but fail at execution?
- Does longer reasoning always improve model accuracy on complex tasks?
- Does reasoning structure match explicit versus implicit task demands?
- Are reasoning models more vulnerable to persuasion than standard models?
- Why do reasoning models fail at learning hidden rules from sparse exceptions?
- Why do smaller models lose reasoning faithfulness more than larger models?
- Can weaker models reliably monitor stronger models during reasoning?
- Why do non-reasoning models work better under extreme decomposition than reasoning models?
- Can explicit optimal algorithms prevent reasoning model collapse at high complexity?
- Do models genuinely reason harder on difficult tasks or just appear to?
- Does fine-tuning push models toward reasoning shortcuts that bypass the chain entirely?
- Why does instruction-following capability decrease as models scale stronger?
- Why does latent reasoning override no-think instructions in models?
- Why does step-by-step reasoning degrade performance on judgment-based tasks?
- Why do models learn reasoning form instead of actual abstract inference?
- Does adding reasoning to models degrade other capabilities like rule inference?
- Why do difficult problems force models to develop reasoning strategies?
- Can models learn when to think versus answer directly?
- Is the reasoning cliff actually a tool-use problem?
- Do reasoning models fail to report processes that actually influence their answers?
- Is reasoning failure caused by task complexity or training distribution gaps?
- Do base models and reasoning models fail in opposite directions on uncertainty?
- Does model scaling improve knowledge storage faster than reasoning ability?
- Can weak models reason better when freed from cognitive load by structure?
- Why do reasoning models fail to improve constrained optimization performance?
- Why might rationales that predict common text patterns fail on hard novel reasoning?
- Can benchmark improvements hide degradation of deliberative reasoning?
- How does scaling reasoning capability actually reduce instruction-following ability?
- Why do simple math problems get worse with longer reasoning chains?
- Can small models solve complex tasks using externalized reasoning graphs?
- Why do models automatically adjust reasoning length to problem difficulty?
- Can models learn when to invoke search during reasoning tasks?
- How does learnability at the observer's current state prevent novelty from breaking model reasoning?
- Are reasoning models more vulnerable to adversarial manipulation than standard models?
- Why do reasoning models fail when input length increases even below context limits?
- Are difficult tasks more monitorable because reasoning externalization becomes necessary?
- Why do strong models struggle more with instruction following than mid-tier ones?
- Why do reasoning-optimized models show no resistance advantage on agreement tasks?
- Can external classifiers reliably decide when a model should reason?
- Does decoupling reasoning from tool use actually improve accuracy?
- How do reasoning-related features behave when trained on near-impossible problems?
- Why do reasoning model failures stem from execution rather than reasoning?
- Why does strategy diversity within reasoning chains improve model generalization?
- Can outcome-focused objectives explain failures in reasoning evaluation?
- Why do models fail on logically equivalent tasks with different data distributions?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- What reward pressure is actually needed to degrade reasoning model monitorability?
- Why do reasoning-optimized models show no sycophancy resistance advantage?
- Why does stronger reasoning reduce model compliance with instructions?
- What changes when reasoning models adopt trajectory-response output formats?
- Can explicit constraint statements override the dominance of surface heuristics?
- What distinguishes coherent reasoning from inaccurate but plausible predictions?
- Do reasoning-enabled and prompt-hardened conditions show the same architectural penalty?
- What limits external scaling when a model lacks reasoning foundation?
- Why does reasoning performance degrade as input length increases?
- Why do reasoning-optimized models still fall for logical fallacies in conversation?
- Why does general reasoning not transfer to knowledge-intensive medical domains?
- Can models distinguish between logical impossibility and their own execution limits?
- Can reasoning models distinguish between new evidence and manipulative reframing?
- Why do reasoning models hide their reliance on hints from evaluators?
- How does task simplification affect analogical reasoning patterns?
- How do models integrate conflicting signals in reasoning tasks?
- Why do reasoning models verbalize reasoning shortcuts less than necessary?
- What makes a background condition relevant to a specific reasoning task?
- Why does target probability matter more than task logical complexity?
- Do reasoning models trade instruction following for deliberative capability?
- What causes reasoning quality to degrade during long research tasks?
- Can a model predict the right action but execute the wrong one?
- Do knowledge access methods like search improve reasoning or just coverage?
- Why do epistemic failure modes cluster around world model limitations?
- Does the heuristic dominance ratio vary predictably across model architectures?
- How do search and reasoning workflows improve forecasting performance over base models?
- How does early commitment in reasoning differ from early exploitation in planning?
- What does pass@k reveal about base model reasoning capacity?
- Can a single SAE feature control reasoning behavior across model families?
- Do reasoning failures stem from strategy or from calculation breakdown?
- How do unstated feasibility constraints affect model decision-making?
- How does reasoning instability prevent models from modeling individuals?
- Why do macro and micro forecasting scales require different reasoning approaches?
- What circuit mechanisms produce belief bias in syllogistic reasoning?
- Why do medical and mathematical tasks require fundamentally different model capabilities?
- What explains the demand effect when available analogies get over-applied?