Line of inquiry
Inquiring lines›What drives capability improvement…›What training and inference approa…›this line of inquiry
What prevents language models from performing systematic logical reasoning?
A broader line of inquiry — a family of 81 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 81
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do reasoning models wander instead of searching systematically?
- Why does extended reasoning fail for search and knowledge retrieval tasks?
- What mechanisms cause reasoning models to wander rather than focus?
- Why do long-context language models struggle with compositional reasoning tasks?
- What sparse mechanistic structures drive reasoning traces in language models?
- What evidence shows that reasoning chains encode token-level functional structure?
- Why do language models struggle with formal logical reasoning and joins?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- Why do language models produce unfaithful chain of thought explanations?
- How do single wrong steps corrupt entire reasoning chains?
- Can long-context models handle compositional reasoning requiring structured logic?
- Can recursive sub-calls decompose reasoning across multiple context chunks?
- Can models compress reasoning chains without external teacher supervision?
- Why do reasoning chains degenerate into undirected exploration at scale?
- Where do humans and language models actually diverge in reasoning ability?
- Can simple structure perturbations reliably expose memorization in reasoning models?
- Does text-only evaluation hide reasoning collapse that tool use could repair?
- How does instance novelty rather than chain length explain reasoning failure?
- When is numeric computation the real bottleneck versus reasoning depth?
- How do recursive language models rethink where to store reasoning?
- Can silent reasoning steps in language models be detected inside the system?
- Why do reasoning models fail on structurally unfamiliar instances?
- How do depth and length constraints emerge in latent reasoning without explicit training?
- Can we distinguish between semantic and symbolic reasoning in language models?
- How do humans and LMs differ on multi-hop reasoning?
- Why do larger reasoning models show cyclicity only in later layers?
- Can cognitive scaffolding replace tool-based reasoning augmentation in language models?
- Can symbolic solvers rescue language models from logical reasoning failures?
- Why does step-by-step reasoning fail when tool outputs get very large?
- Why do expert reasoners skip steps that novices must state explicitly?
- What makes deterministic recursive reasoning models underperform on multi-solution tasks?
- Does sentence-level granularity capture enough structure for complex reasoning tasks?
- Can reasoning evaluation metrics reward actual reasoning instead of theater?
- Why do some reasoning models fail to detect redundancy in concurrent coordination?
- Does architectural design matter more than model scale for reasoning tasks?
- Can single-hop knowledge automatically compose into multi-hop capability?
- Can a tiny recursive network beat billion-parameter models on hard problems?
- Why do open-source models trained on proprietary outputs still fail at reasoning?
- Can reasoning benchmarks separate logic from believability?
- Why do language models struggle with backward reasoning compared to forward?
- What design changes could make constraint inference more reliable without explicit cuing?
- Does reasoning efficiency transfer to tasks without ground truth dependency graphs?
- Can surface heuristics override implicit constraints in domain-specific reasoning?
- How can entailment benchmarks separate genuine reasoning from memorization effects?
- What makes deductive reasoning so brittle in language models overall?
- Why does the Chinese Room argument miss the deeper abstraction problem?
- Why does second-hop reasoning fail when composed with out-of-distribution triples?
- Why does removing semantic content collapse reasoning in language models?
- How does recombining partial trajectories maintain coherence in natural language reasoning?
- How does SONAR embedding quality affect downstream reasoning accuracy?
- How does the frame problem differ between symbolic and statistical reasoning systems?
- Why are pairwise relations insufficient for representing higher-order multi-hop reasoning?
- Why does comparison reasoning generalize better than composition reasoning?
- Why do language models produce verbose reasoning when asked to think step by step?
- How do alternative hypothesis checks reduce confirmation bias in code reasoning?
- Do tool-enabled reasoning models close the gap on constraint satisfaction?
- How do deterministic symbolic solvers improve the reliability of language model reasoning?
- Do sparse arithmetic circuits explain all language model reasoning abilities?
- What architectural properties of deterministic models block multi-solution reasoning?
- Which RAG sub-decisions are actually pattern matching versus reasoning intensive?
- Why does output alignment fail to catch internally incoherent reasoning?
- What failure modes emerge when scheme classification feeds downstream reasoning pipelines?
- Can reasoning in free text then formatting separately recover performance?
- What causes snowball errors to accumulate across reasoning steps in language models?
- How do dependency errors propagate through incorrectly formalized definitions?
- How do causal chains enforce long-horizon length differently than instruction-based tasks?
- How does inductive reasoning from partial evidence enable hypothesis formation?
- How does evidence retrieval affect compositional reasoning in language models?
- What makes constraint satisfaction problems epistemically cleaner than other reasoning tasks?
- What implicit premises do language models skip even with correct surface reasoning?
- How do failed branches remain in context and contaminate subsequent reasoning?
- What is the mechanistic signature when models chain facts never presented together?
- How does open-ended evolver reasoning identify patterns across heterogeneous user trajectories?
- Can long-context readers handle compositional tasks or just semantic search?
- Why do sparse per-step errors accumulate undetected across delegated tasks?
- What neuroscience evidence suggests language networks are not optimized for reasoning?
- How does contrapositive augmentation change the tractability of reasoning tasks?
- Do distributed relational tasks consistently underperform local classification across NLP domains?
- What makes Compound-QA expose weaknesses in monologue reasoning?
- Why do unresolved items cluster in structured patterns rather than randomly?
- How much does schema bloat actually degrade reasoning in large language models?