Theme of inquiry
How do different training strategies affect reasoning and generalization?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
49 specific questions
- Can single-problem fine-tuning match full RL pipeline reasoning gains?
- Does format-based pretraining determine how models respond to reinforcement learning?
- Why does RL improve sampling efficiency but not expand capability boundaries?
- Why do overtrained domains show different RL training outcomes than novel tasks?
- What makes pretraining composition more important than reward engineering?
- How do sparse parameter updates enable when-not-how training to work?
- How does baseline capability level affect RL improvement ceiling?
46 specific questions
- Does RL teach models when to use reasoning or how to reason?
- Does RL primarily teach when to use reasoning or how to reason?
- Can RL create new reasoning primitives that pretraining never established?
- When does reinforcement learning actually produce true reasoning gains in models?
- Does RL amplify existing reasoning or create genuinely new computational strategies?
- Can extended RL training unlock genuinely new reasoning strategies models cannot discover otherwise?
- How does RL refine reasoning paths without simply adding model capability?
31 specific questions
- What causes policy entropy collapse in reasoning-focused reinforcement learning?
- Does policy entropy collapse limit how many iterations of reasoning training work?
- What happens to model reasoning when policy entropy collapses during RL?
- Why does policy entropy collapse limit reasoning and dialogue RL scaling?
- Does policy entropy collapse prevent inference-time search from finding solutions?
- Why does policy entropy collapse when scaling RL for reasoning?
- How does policy entropy during training affect search discipline during inference?
8 specific questions
- How do continuous concept tokens explore multiple reasoning paths without explicit sampling?
- How do soft thinking and token-level mixtures explore multiple paths simultaneously?
- How do soft token mixtures enable parallel reasoning exploration without explicit training?
- How does continuous soft thinking explore multiple paths without explicit training?
- How do continuous concept tokens compare to latent trajectory sampling?
- How does soft thinking compare to sampling multiple independent reasoning paths?
- How does soft thinking achieve stochastic exploration without explicit training?
20 specific questions
- Can explicitly optimizing for semantic diversity during RL training improve both quality and variation?
- Does semantic diversity in output space compete with reward-component diversity?
- Does optimizing directly for semantic diversity improve both reasoning quality and exploration?
- Why does outcome-based RL specifically lose diversity during training?
- How does forced exploration through diversity rewards differ from suppression-based negative reinforcement?
- Why does exploration diversity behave differently under reinforcement learning versus supervised fine-tuning?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?