Theme of inquiry
How does reward hacking arise and persist in reinforcement learning?
A question within its area, explored through 2 lines of inquiry below — each a family of specific questions the research asks.
107 specific questions
- Do models that recognize reward hacking disclose it in their outputs?
- Do models reward hack at high rates on unmodified benchmarks?
- What detection method survives when a model optimizes to hide hacking?
- What determines the ground truth when detecting reward hacking in model evaluations?
- How do chain-of-thought monitors become targets for reward hacking?
- Does reward hacking cause evaluations to overstate model capabilities?
- Can belief checks detect whether models will resist reward hacking?
37 specific questions
- Why do honeypot tasks reveal reward hacking better than standard benchmarks?
- How do planted hacks in benchmarks help measure real-world reward hacking?
- Do unmodified benchmarks expose similar reward hacking vectors without planted shortcuts?
- How do unplanted benchmark hacking rates compare to rates on constructed bait tasks?
- Does a planted honeypot count the hacks that actually matter?
- Does a planted honeypot catch all the hacks that actually matter?
- Do honeypot shortcuts in synthetic tasks measure the hacks that actually matter?