Line of inquiry
Inquiring lines›How can we optimize language model…›How do test-time resources and tra…›this line of inquiry
What causes reasoning models to fail on structurally novel problems?
A broader line of inquiry — a family of 78 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 78
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do reasoning models wander instead of searching systematically?
- What mechanisms cause reasoning models to wander rather than focus?
- How does instance novelty rather than chain length explain reasoning failure?
- Do reasoning models switch approaches when encountering local difficulty?
- Why do reasoning models fail at learning hidden rules from sparse exceptions?
- When does explicit reasoning actually degrade performance on a task?
- Why do reasoning models fail on structurally unfamiliar instances?
- Can explicit optimal algorithms prevent reasoning model collapse at high complexity?
- Is reasoning failure caused by task complexity or training distribution gaps?
- Where do humans and language models actually diverge in reasoning ability?
- Can reasoning models succeed at logic but fail at execution?
- Why does extended reasoning fail for search and knowledge retrieval tasks?
- Why do smaller models lose reasoning faithfulness more than larger models?
- How can high benchmark performance mask broken reasoning in AI systems?
- Why do non-reasoning models work better under extreme decomposition than reasoning models?
- When is numeric computation the real bottleneck versus reasoning depth?
- Why do reasoning chains degenerate into undirected exploration at scale?
- How do single wrong steps corrupt entire reasoning chains?
- Does text-only evaluation hide reasoning collapse that tool use could repair?
- Why do longer reasoning chains explore like tourists instead of scientists?
- Why might rationales that predict common text patterns fail on hard novel reasoning?
- Why do language models struggle with formal logical reasoning and joins?
- What makes deterministic recursive reasoning models underperform on multi-solution tasks?
- Why do reasoning models fail to improve constrained optimization performance?
- Is the reasoning cliff actually a tool-use problem?
- Can benchmark improvements hide degradation of deliberative reasoning?
- Why does step-by-step reasoning fail when tool outputs get very large?
- Why do language models struggle with backward reasoning compared to forward?
- Why do models automatically adjust reasoning length to problem difficulty?
- Why do reasoning models fail when input length increases even below context limits?
- What makes diverse reasoning sources more valuable than deeper single paths?
- Why do models skip steps that would make reasoning clearer?
- Do base models and reasoning models fail in opposite directions on uncertainty?
- Why do expert reasoners skip steps that novices must state explicitly?
- Why does per-step deliberation lose global perspective compared to dynamic discovery?
- Can symbolic solvers rescue language models from logical reasoning failures?
- Why do language model reasoning chains look fluent when they deviate from the task?
- Can outcome-focused objectives explain failures in reasoning evaluation?
- Why do reasoning-optimized models still fall for logical fallacies in conversation?
- Why do open-source models trained on proprietary outputs still fail at reasoning?
- Can reasoning benchmarks separate logic from believability?
- How much reasoning depth do we actually need for most real-world tasks?
- Why do some reasoning models fail to detect redundancy in concurrent coordination?
- Why do reasoning model failures stem from execution rather than reasoning?
- What causes reasoning quality to degrade during long research tasks?
- How do reasoning-related features behave when trained on near-impossible problems?
- Why does removing semantic content collapse reasoning in language models?
- Why does the Chinese Room argument miss the deeper abstraction problem?
- How does active reasoning through interaction differ from passive single-turn problem solving?
- Why does comparison reasoning generalize better than composition reasoning?
- What makes a background condition relevant to a specific reasoning task?
- Can verifier-guided search catch factual errors that reasoning training cannot?
- What failure modes emerge when scheme classification feeds downstream reasoning pipelines?
- What limits external scaling when a model lacks reasoning foundation?
- Why does more inference compute amplify wandering rather than solving it?
- How does the frame problem differ between symbolic and statistical reasoning systems?
- Why do contrastive reasoning approaches outperform single-path belief evaluation?
- Can models distinguish between logical impossibility and their own execution limits?
- How do failed branches remain in context and contaminate subsequent reasoning?
- Do reasoning failures stem from strategy or from calculation breakdown?
- Why does output alignment fail to catch internally incoherent reasoning?
- Why do epistemic failure modes cluster around world model limitations?
- How does early commitment in reasoning differ from early exploitation in planning?
- What architectural properties of deterministic models block multi-solution reasoning?
- What causes snowball errors to accumulate across reasoning steps in language models?
- How can minimal pairs expose reasoning failures that single-instance accuracy metrics miss?
- Do sparse arithmetic circuits explain all language model reasoning abilities?
- What makes constraint satisfaction problems epistemically cleaner than other reasoning tasks?
- How does reasoning instability prevent models from modeling individuals?
- What distinguishes systematic search from wandering exploration in reasoning?
- What neuroscience evidence suggests language networks are not optimized for reasoning?
- How do semantic failure modes map to attentional and intentional layers?
- What happens to iterative search quality when reasoning depth is unconstrained?
- Why does naive randomness fail to improve stochastic latent reasoning models?
- Why does scheme classification require more cognitive load than identifying premises?
- What makes Compound-QA expose weaknesses in monologue reasoning?
- Why does single-shot learning fail in REVTHINK's multi-source reasoning tasks?
- What tree depth is achievable before GPU memory becomes the bottleneck?