Theme of inquiry

How do different training strategies affect reasoning and generalization?

A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.


What pretraining choices and baseline capability constrain reinforcement learning gains?

49 specific questions

See all 49 questions in this line of inquiry
Does reinforcement learning teach reasoning or just when to reason?

46 specific questions

See all 46 questions in this line of inquiry
How does policy entropy collapse constrain reasoning-focused reinforcement learning?

31 specific questions

See all 31 questions in this line of inquiry
How do soft continuous representations explore multiple reasoning paths simultaneously?

8 specific questions

See all 8 questions in this line of inquiry
Why does reinforcement learning suppress output diversity compared to supervised fine-tuning?

20 specific questions

See all 20 questions in this line of inquiry