Theme of inquiry

How do different reward signals and mechanisms drive agent learning?

A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.


How do spurious versus genuine rewards shape model reasoning and behavior?

78 specific questions

See all 78 questions in this line of inquiry
Can inoculation prompting prevent emergent misalignment after reward hacking?

33 specific questions

See all 33 questions in this line of inquiry
How do pretraining biases affect reward signal effectiveness in RLVR?

91 specific questions

See all 91 questions in this line of inquiry
What makes step-level supervision effective for complex reasoning traces?

49 specific questions

See all 49 questions in this line of inquiry
How can reward models capture diverse human preferences without excluding minority populations?

48 specific questions

See all 48 questions in this line of inquiry