Line of inquiry
Inquiring lines›How should agents manage and coord…›How effectively can inference-time…›this line of inquiry
Why do reasoning models fail at systematic problem-solving and search?
A broader line of inquiry — a family of 88 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 88
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do reasoning models wander instead of searching systematically?
- What mechanisms cause reasoning models to wander rather than focus?
- What sparse mechanistic structures drive reasoning traces in language models?
- Why does extended reasoning fail for search and knowledge retrieval tasks?
- Why do reasoning models fail at learning hidden rules from sparse exceptions?
- What evidence shows that reasoning chains encode token-level functional structure?
- Why do smaller models lose reasoning faithfulness more than larger models?
- Can models compress reasoning chains without external teacher supervision?
- Why do models overthink underspecified problems instead of rejecting them?
- Can simple structure perturbations reliably expose memorization in reasoning models?
- How do single wrong steps corrupt entire reasoning chains?
- Why do reasoning chains degenerate into undirected exploration at scale?
- Can recursive sub-calls decompose reasoning across multiple context chunks?
- Why do long-context language models struggle with compositional reasoning tasks?
- Why do language models produce unfaithful chain of thought explanations?
- How does instance novelty rather than chain length explain reasoning failure?
- Where do humans and language models actually diverge in reasoning ability?
- Why do language models struggle with formal logical reasoning and joins?
- Does text-only evaluation hide reasoning collapse that tool use could repair?
- How do recursive language models rethink where to store reasoning?
- Can language models reason without relying on surface level pattern matching?
- Why do reasoning models fail on structurally unfamiliar instances?
- Why do larger reasoning models show cyclicity only in later layers?
- Can long-context models handle compositional reasoning requiring structured logic?
- Why do language models generate reasoning tokens after internally deciding the answer?
- Can we distinguish between semantic and symbolic reasoning in language models?
- When is numeric computation the real bottleneck versus reasoning depth?
- Can cognitive scaffolding replace tool-based reasoning augmentation in language models?
- Is reasoning failure caused by task complexity or training distribution gaps?
- How do humans and LMs differ on multi-hop reasoning?
- Can layer-wise prediction stabilization identify when genuine reasoning has stopped?
- Can explicit optimal algorithms prevent reasoning model collapse at high complexity?
- Can reasoning evaluation metrics reward actual reasoning instead of theater?
- Why do expert reasoners skip steps that novices must state explicitly?
- Does more thinking always help large language models or sometimes hurt?
- Why does step-by-step reasoning fail when tool outputs get very large?
- What makes diverse reasoning sources more valuable than deeper single paths?
- Can activation patching reveal which reasoning steps actually matter?
- Why does the first generated token trigger collapse of task superposition?
- Does sentence-level granularity capture enough structure for complex reasoning tasks?
- Can small models solve complex tasks using externalized reasoning graphs?
- When should a system decide to retrieve versus reason alone?
- Can inserted errors in reasoning drafts produce predictable downstream effects?
- Can a tiny recursive network beat billion-parameter models on hard problems?
- What makes deterministic recursive reasoning models underperform on multi-solution tasks?
- Why do some reasoning models fail to detect redundancy in concurrent coordination?
- How does Self-Discover compare to the cognitive tools approach?
- Can marginal hints integrate better into reasoning than comprehensive explanations?
- Why do language models struggle with backward reasoning compared to forward?
- Why do language models fail at grounding and inference?
- What distinguishes the convergence patterns between reasoning and lexical variation tasks?
- Does reasoning efficiency transfer to tasks without ground truth dependency graphs?
- Can external classifiers reliably decide when a model should reason?
- Can scaffolding frameworks isolate inductive reasoning from deductive confounds?
- What design changes could make constraint inference more reliable without explicit cuing?
- Can instance-adaptive reasoning happen without sequential token dependencies?
- Why does the Chinese Room argument miss the deeper abstraction problem?
- What quality filters distinguish useful reasoning enrichment from shallow repetition?
- Why does second-hop reasoning fail when composed with out-of-distribution triples?
- How does recombining partial trajectories maintain coherence in natural language reasoning?
- How does SONAR embedding quality affect downstream reasoning accuracy?
- Can bounded workspaces prevent overthinking better than summarization alone?
- What makes a background condition relevant to a specific reasoning task?
- Why do language models produce verbose reasoning when asked to think step by step?
- Why do reasoning-optimized models still fall for logical fallacies in conversation?
- What failure modes emerge when scheme classification feeds downstream reasoning pipelines?
- Why are pairwise relations insufficient for representing higher-order multi-hop reasoning?
- How does collaboration itself become a degradation mechanism in reasoning tasks?
- What causes snowball errors to accumulate across reasoning steps in language models?
- Do sparse arithmetic circuits explain all language model reasoning abilities?
- Can dataset design systematically expand reasoning graph diameter?
- How does inductive reasoning from partial evidence enable hypothesis formation?
- Why does output alignment fail to catch internally incoherent reasoning?
- Can reasoning in free text then formatting separately recover performance?
- What architectural properties of deterministic models block multi-solution reasoning?
- How do humans and R1 models differ in information gain patterns?
- How do dependency errors propagate through incorrectly formalized definitions?
- What makes multi-turn critique trajectories more effective than single-turn reasoning chains?
- What neuroscience evidence suggests language networks are not optimized for reasoning?
- What is the mechanistic signature when models chain facts never presented together?
- What implicit premises do language models skip even with correct surface reasoning?
- How does open-ended evolver reasoning identify patterns across heterogeneous user trajectories?
- Why does naive randomness fail to improve stochastic latent reasoning models?
- Why do smaller LLMs fail at zero-shot argument scheme classification?
- Why does scheme classification require more cognitive load than identifying premises?
- Why does uniform averaging across all tokens dilute the reasoning signal?
- What makes Compound-QA expose weaknesses in monologue reasoning?
- Why does the second loop do most of the productive refinement work?