Theme of inquiry
What training and inference approaches improve model reasoning capability?
A question within its area, explored through 8 lines of inquiry below — each a family of specific questions the research asks.
82 specific questions
- Can models maintain reasoning-output coupling while improving domain accuracy?
- Does supervised fine-tuning improve accuracy while damaging the quality of reasoning?
- Can fine-tuning ever teach semantic inference instead of amplifying training shortcuts?
- Why does instruction tuning hurt knowledge-intensive tasks more than reasoning tasks?
- Does task diversity in pretraining data transfer reasoning better than larger models?
- How does data quality mismatch create reasoning degradation in supervised fine-tuning?
- Does domain training degrade reasoning ability even when benchmark scores rise?
34 specific questions
- Why do deliberately corrupted reasoning traces sometimes generalize better than correct ones?
- Why do corrupted traces maintain performance as well as correct traces?
- Do corrupted reasoning traces teach something different than pure success traces?
- Why do corrupted reasoning traces sometimes generalize better than correct ones?
- What makes some reasoning traces better supervision than others despite equal accuracy?
- Why are incorrect reasoning traces longer than correct ones?
- Can training on reasoning traces teach actual self-correction or only confident first answers?
81 specific questions
- Why do reasoning models wander instead of searching systematically?
- Why does extended reasoning fail for search and knowledge retrieval tasks?
- What mechanisms cause reasoning models to wander rather than focus?
- Why do long-context language models struggle with compositional reasoning tasks?
- What sparse mechanistic structures drive reasoning traces in language models?
- What evidence shows that reasoning chains encode token-level functional structure?
- Why do language models struggle with formal logical reasoning and joins?
41 specific questions
- What causes policy entropy collapse in reasoning-focused reinforcement learning?
- Does policy entropy collapse limit how many iterations of reasoning training work?
- What happens to model reasoning when policy entropy collapses during RL?
- Why does policy entropy collapse limit reasoning and dialogue RL scaling?
- Does policy entropy collapse prevent inference-time search from finding solutions?
- Why does policy entropy collapse when scaling RL for reasoning?
- How does policy entropy during training affect search discipline during inference?
21 specific questions
- Why does training data format shape reasoning strategy more than domain content?
- How much does training data format influence reasoning strategy versus domain content?
- How does training format shape reasoning strategy more than content?
- Does training data format shape reasoning strategy more than domain content?
- Why does training data format shape reasoning strategy more than content?
- Does training data format shape model reasoning more than domain content?
- Can training format itself shape what reasoning strategy a model learns?
100 specific questions
- Do base models contain latent reasoning that minimal training can unlock?
- Can minimal training signals unlock latent reasoning capability in base models?
- What latent reasoning capability do base models already possess before training?
- Why do reasoning gains resist clear attribution to specific training changes?
- Can models reason at inference without specialized internal training?
- Can models possess latent reasoning capability that training signals fail to unlock?
- Can minimal training signals unlock reasoning already latent in pretrained representations?
102 specific questions
- Can minimal reasoning steps match verbose reasoning accuracy?
- Why might latent reasoning capture types of thinking that verbalized CoT cannot?
- Can steering a single latent feature replicate chain-of-thought performance?
- How does extended thinking affect variance in reasoning model outputs?
- When does explicit reasoning actually degrade performance on a task?
- Can latent reasoning scale test-time compute without verbalized tokens or special training?
- Can latent reasoning in continuous space scale beyond supervised reasoning tasks?
79 specific questions
- Why does augmenting symbolic reasoning outperform replacing it entirely?
- Does structured decomposition improve LLM reasoning in other compound tasks?
- What makes structural logic correlate so strongly with contextual consistency?
- Can reasoning chains work without logical validity?
- Why do format and structure matter more than actual content in reasoning?
- Can symbolic solvers reliably replace LLM reasoning for logical tasks?
- Do reasoning systems reuse cognitive structures across unrelated topics?