Theme of inquiry
How effectively can inference-time compute improve model reasoning performance?
A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.
88 specific questions
- Why do reasoning models wander instead of searching systematically?
- What mechanisms cause reasoning models to wander rather than focus?
- What sparse mechanistic structures drive reasoning traces in language models?
- Why does extended reasoning fail for search and knowledge retrieval tasks?
- Why do reasoning models fail at learning hidden rules from sparse exceptions?
- What evidence shows that reasoning chains encode token-level functional structure?
- Why do smaller models lose reasoning faithfulness more than larger models?
33 specific questions
- What latent reasoning capability do base models already possess before training?
- Do base models contain latent reasoning that minimal training can unlock?
- What mechanisms activate latent reasoning capabilities already present in base models?
- Can minimal training signals unlock latent reasoning capability in base models?
- Can models possess latent reasoning capability that training signals fail to unlock?
- Does the base model already contain latent reasoning capability?
- Does latent reasoning capability exist in base models before any training?
25 specific questions
- Does supervised fine-tuning improve accuracy while damaging the quality of reasoning?
- Does SFT degrade reasoning quality while improving domain accuracy?
- Does supervised fine-tuning improve reasoning or just response formatting?
- Why does fine-tuning degrade reasoning quality even as accuracy improves?
- Why does domain accuracy improve while reasoning quality degrades after supervised fine-tuning?
- Why does SFT reduce reasoning quality even when improving domain accuracy?
- How does supervised fine-tuning degrade chain-of-thought faithfulness over time?
46 specific questions
- What makes multi-paradigm chaining a distinct reasoning topology?
- Why does reasoning graph topology evolve differently across training phases?
- Can graph cyclicity and topology predict when reasoning systems achieve breakthrough insights?
- Can removing hierarchy from dual-recurrence models improve reasoning performance?
- How do graph-based reasoning topologies map to multi-agent interaction patterns?
- How do graph topology properties like cyclicity and diameter affect reasoning quality?
- Why do different reasoning chains surface different relevant facts?
33 specific questions
- Do tool-enabled reasoning models close the gap on constraint satisfaction?
- Can symbolic solvers reliably replace LLM reasoning for logical tasks?
- Can symbolic solvers rescue language models from logical reasoning failures?
- How do deterministic symbolic solvers improve the reliability of language model reasoning?
- Can tools unlock reasoning strategies that require abstract insight beyond computation?
- Can external verifiers replace reasoning trace quality in solution guarantees?
- Why do semi-formal templates improve verification accuracy over unstructured reasoning?
29 specific questions
- Why do knowledge and reasoning train in different network layers?
- Why do higher network layers capture procedural knowledge but lower layers store facts?
- What separates knowledge from reasoning in neural network layers?
- How do retrieval heads interact with layer-level separation of knowledge and reasoning?
- What is the difference between procedural knowledge and factual retrieval in reasoning?
- How do knowledge layers differ functionally from reasoning layers in networks?
- How do procedural versus factual knowledge differ in pretraining versus fine-tuning?