Theme of inquiry

Why do models pursue reward hacking despite training constraints?

A question within its area, explored through 3 lines of inquiry below — each a family of specific questions the research asks.


How does optimization for reward create emergent misalignment in language models?

75 specific questions

See all 75 questions in this line of inquiry
How can evaluations be made robust against model reward hacking?

53 specific questions

See all 53 questions in this line of inquiry
What evaluation methods best detect reward hacking in AI agents?

51 specific questions

See all 51 questions in this line of inquiry