Theme of inquiry

How do reward models and preferences optimize agent training?

A question within its area, explored through 4 lines of inquiry below — each a family of specific questions the research asks.


Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones?

100 specific questions

See all 100 questions in this line of inquiry
Can iterative DPO substitute for online RL in studying misalignment?

31 specific questions

See all 31 questions in this line of inquiry
How do reward signal properties affect model reasoning and safety?

136 specific questions

See all 136 questions in this line of inquiry
What makes process supervision effective for training complex reasoning models?

57 specific questions

See all 57 questions in this line of inquiry