Theme of inquiry
How do different reward signals and mechanisms drive agent learning?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
78 specific questions
- Do spurious rewards activate reasoning without teaching new skills?
- Can decomposing rewards into prompt-free and prompt-related components fix this blindspot?
- Can structured rewards still teach models when spurious rewards also work?
- Why do spurious reward signals improve reasoning for some pretrained models?
- Can random rewards improve reasoning models if pretraining is suitable?
- How do reward models benefit from extended thinking during evaluation scoring?
- How can reward-seeking remain hidden when graders reward the intended behavior?
33 specific questions
- Can inoculation prompting prevent emergent misalignment after reward hacking in production RL?
- Can inoculation prompting reduce alignment faking by reframing reward hacking as acceptable?
- Why does inoculation prompting prevent misaligned generalization from reward hacking?
- What mechanism drives emergent misalignment in reward-hacked models instead?
- Does reward-seeking mediate emergent misalignment after reward hacking?
- Does inoculation prompting at training time reduce which reward hacks generalize?
- Does this misalignment pattern appear outside reward hacking environments?
91 specific questions
- Does careful reward engineering matter if pretraining determines RLVR effectiveness?
- What makes reward signal sources substitutable across verifier-free RL patterns?
- Is elaborate reward shaping necessary if the pretrained prior already contains good solutions?
- How do reward signals in RLVR interact with pretraining biases?
- What makes current learned reward models fail across different domains?
- Can models trained with RL on pretraining data avoid reward hacking seen in RLHF?
- What other downstream metrics could serve as RL reward sources?
49 specific questions
- What does process supervision reveal about step-level reasoning versus outcome rewards?
- Do process reward models need different supervision strategies by domain?
- What makes process-level supervision better than outcome-only reward signals?
- Can process reward models work on branching reasoning traces with backtracking?
- How does process supervision relate to execution-signaled feedback approaches?
- How can process reward models handle branching and revisiting in reasoning traces?
- What are the actual limits of sibling comparison versus trained process reward models?
48 specific questions
- How do aggregate reward models fail to capture minority user preferences?
- How do aggregate reward models systematically exclude minority preferences?
- Can reward models distinguish between personal preference and community consensus?
- How do aggregate reward models systematically exclude minority perspectives?
- Why does single-reward RLHF fail to represent diverse human preferences?
- Can a single AI judge capture diverse human preferences or does it collapse them?
- What explicit safeguards should limit personalization in deployed reward models?