Theme of inquiry

How does reward hacking arise and persist in reinforcement learning?

A question within its area, explored through 2 lines of inquiry below — each a family of specific questions the research asks.


Can we reliably detect when models game evaluations?

107 specific questions

See all 107 questions in this line of inquiry
Do honeypot benchmarks validly measure reward hacking better than standard tests?

37 specific questions

See all 37 questions in this line of inquiry