Theme of inquiry
How do test-time resources and training improve model reasoning?
A question within its area, explored through 13 lines of inquiry below — each a family of specific questions the research asks.
45 specific questions
- Why might latent reasoning capture types of thinking that verbalized CoT cannot?
- Can latent reasoning scale test-time compute without verbalized tokens or special training?
- Can latent reasoning achieve the same substitution without tokens?
- Can continuous latent reasoning match discrete chain-of-thought without training modifications?
- Can latent reasoning scale test-time compute without verbal tokens?
- Can reasoning happen in latent space without chain of thought?
- Can latent reasoning in continuous space scale beyond supervised reasoning tasks?
21 specific questions
- Why does training data format shape reasoning strategy more than domain content?
- How much does training data format influence reasoning strategy versus domain content?
- How does training format shape reasoning strategy more than content?
- Does training data format shape reasoning strategy more than domain content?
- Why does training data format shape reasoning strategy more than content?
- Does training data format shape model reasoning more than domain content?
- Can training format itself shape what reasoning strategy a model learns?
64 specific questions
- Do base models contain latent reasoning that minimal training can unlock?
- Can minimal training signals unlock latent reasoning capability in base models?
- What latent reasoning capability do base models already possess before training?
- Can models possess latent reasoning capability that training signals fail to unlock?
- Can minimal training signals unlock reasoning already latent in pretrained representations?
- Does the base model already contain latent reasoning capability?
- What mechanisms activate latent reasoning capabilities already present in base models?
64 specific questions
- Does scaling reasoning capability create tradeoffs with instruction following?
- Why do models learn reasoning form instead of actual abstract inference?
- Why do instruction following and reasoning capability trade off in training?
- Does reasoning structure match explicit versus implicit task demands?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- Why does instruction-following capability decrease as models scale stronger?
- Can models learn when to think versus answer directly?
27 specific questions
- Can reasoning learned from language modeling actually transfer to knowledge-intensive domains?
- Does model scaling improve knowledge storage faster than reasoning ability?
- Why does reasoning training improve math but hurt knowledge tasks?
- Can mathematical reasoning improvements transfer across problem subdomains?
- What makes knowledge-rich specialized domains structurally different from general reasoning tasks?
- Does task diversity in pretraining data transfer reasoning better than larger models?
- Why does general reasoning not transfer to knowledge-intensive medical domains?
59 specific questions
- Can reasoning traces prove models are actually reasoning versus mimicking?
- Why do reasoning traces fail to accurately reflect model decision-making?
- Can reasoning traces serve purposes beyond producing the final answer itself?
- Why do corrupted traces maintain performance as well as correct traces?
- How do reasoning traces fail to represent what models actually computed?
- What makes a reasoning trace causally sufficient versus merely stylistically plausible?
- Can corrupted reasoning traces be reliably distinguished from correct ones?
29 specific questions
- What causes policy entropy collapse in reasoning-focused reinforcement learning?
- Does policy entropy collapse limit how many iterations of reasoning training work?
- What happens to model reasoning when policy entropy collapses during RL?
- Why does policy entropy collapse limit reasoning and dialogue RL scaling?
- Does policy entropy collapse prevent inference-time search from finding solutions?
- Why does policy entropy collapse when scaling RL for reasoning?
- How does policy entropy collapse constrain token-level distribution in reasoning?
36 specific questions
- Why do language models substitute parametric knowledge over retrieved context mid-reasoning?
- Can models internalize retrieved context as static parametric knowledge?
- Can models recover knowledge with completely unrelated retraining tasks?
- Does finetuning facts into weights overwrite existing model capabilities?
- How does parametric knowledge sabotage context-grounded question answering?
- What makes some contexts learnable as rules versus requiring model retraining?
- How do we distinguish knowledge encoding from knowledge usage in models?
78 specific questions
- Why do reasoning models wander instead of searching systematically?
- What mechanisms cause reasoning models to wander rather than focus?
- How does instance novelty rather than chain length explain reasoning failure?
- Do reasoning models switch approaches when encountering local difficulty?
- Why do reasoning models fail at learning hidden rules from sparse exceptions?
- When does explicit reasoning actually degrade performance on a task?
- Why do reasoning models fail on structurally unfamiliar instances?
38 specific questions
- Why does fine-tuning degrade reasoning quality even as accuracy improves?
- Does supervised fine-tuning improve accuracy while damaging the quality of reasoning?
- Does SFT degrade reasoning quality while improving domain accuracy?
- How does data quality mismatch create reasoning degradation in supervised fine-tuning?
- Can fine-tuning ever teach semantic inference instead of amplifying training shortcuts?
- Why does supervised fine-tuning degrade reasoning quality despite raising accuracy?
- Does supervised fine-tuning improve reasoning or just response formatting?
66 specific questions
- Does structured decomposition improve LLM reasoning in other compound tasks?
- Can language models reason without relying on learned semantic patterns?
- Does more thinking always help large language models or sometimes hurt?
- Why do language models imitate reasoning form without abstract inference capability?
- Can language models reason without relying on surface level pattern matching?
- Do language models build world models or just task-specific heuristics?
- Why does augmenting symbolic reasoning outperform replacing it entirely?
47 specific questions
- What evidence shows that reasoning chains encode token-level functional structure?
- What makes multi-paradigm chaining a distinct reasoning topology?
- Can small models solve complex tasks using externalized reasoning graphs?
- Can single-hop knowledge automatically compose into multi-hop capability?
- Can recursive sub-calls decompose reasoning across multiple context chunks?
- How do humans and LMs differ on multi-hop reasoning?
- Can long-context models handle compositional reasoning requiring structured logic?
50 specific questions
- Do models deliberately hide influences from their reasoning traces?
- Are reasoning models more vulnerable to persuasion than standard models?
- Can activation probes detect reasoning that models omit from text?
- Are reasoning models more vulnerable to adversarial manipulation than standard models?
- Why does latent reasoning override no-think instructions in models?
- How do sycophancy hints stay invisible despite appearing in reasoning chains?
- Do reasoning models fail to report processes that actually influence their answers?