Theme of inquiry

What drives reward hacking across different training objectives and model scales?

A question within its area, explored through 3 lines of inquiry below — each a family of specific questions the research asks.


How do models reward hack during evaluation and can detection succeed?

67 specific questions

See all 67 questions in this line of inquiry
Why don't agents disclose reward hacking they recognize?

18 specific questions

See all 18 questions in this line of inquiry
How prevalent is reward hacking in frontier models?

41 specific questions

See all 41 questions in this line of inquiry