Theme of inquiry
How do reward models and preferences optimize agent training?
A question within its area, explored through 4 lines of inquiry below — each a family of specific questions the research asks.
100 specific questions
- Does RL amplify existing reasoning or create genuinely new computational strategies?
- Can reinforcement learning add new capabilities or only remove inaccurate knowledge?
- How does RL refine reasoning paths without simply adding model capability?
- Does RL primarily teach when to use reasoning or how to reason?
- Can RL create new reasoning primitives that pretraining never established?
- When does RL discover genuinely novel reasoning strategies versus timing optimization?
- When does reinforcement learning actually produce true reasoning gains in models?
31 specific questions
- Does iterative DPO preserve the same misalignment mechanisms as online reinforcement learning?
- How does iterative DPO compare to full RLVR for studying emergent misalignment?
- Does iterative DPO generalize like online reinforcement learning?
- Does iterative DPO generalize identically to online RL on the same benchmark tasks?
- Can iterative DPO serve as a tractable proxy for studying on-policy misalignment?
- What misaligned behaviors does iterative DPO produce compared to online RL?
- What theoretical argument connects iterative DPO dynamics to online RL learning?
136 specific questions
- Can structured rewards still teach models when spurious rewards also work?
- Do spurious rewards activate reasoning without teaching new skills?
- Why do spurious reward signals improve reasoning for some pretrained models?
- Can models exploit reward systems while appearing to follow safety instructions?
- Does careful reward engineering matter if pretraining determines RLVR effectiveness?
- Can decomposing rewards into prompt-free and prompt-related components fix this blindspot?
- Can random rewards improve reasoning models if pretraining is suitable?
57 specific questions
- What does process supervision reveal about step-level reasoning versus outcome rewards?
- Do process reward models need different supervision strategies by domain?
- What makes process-level supervision better than outcome-only reward signals?
- Can process reward models work on branching reasoning traces with backtracking?
- How does process supervision relate to execution-signaled feedback approaches?
- How can process reward models handle branching and revisiting in reasoning traces?
- What are the actual limits of sibling comparison versus trained process reward models?