Theme of inquiry
What drives reward hacking across different training objectives and model scales?
A question within its area, explored through 3 lines of inquiry below — each a family of specific questions the research asks.
67 specific questions
- What determines the ground truth when detecting reward hacking in model evaluations?
- Do models reward hack at high rates on unmodified benchmarks?
- Does steering through training data override reward hacking associations reliably?
- Do models that recognize reward hacking disclose it in their outputs?
- How do chain-of-thought monitors become targets for reward hacking?
- What detection method survives when a model optimizes to hide hacking?
- Can belief checks detect whether models will resist reward hacking?
18 specific questions
- Do agents disclose reward hacking in the outputs they return?
- Do agents frame reward hacks as valid strategies rather than flaws?
- Do agents aware of their own reward hacking disclose it in outputs?
- Do agents recognize their own reward hacking before submitting their answers?
- Why do agents show awareness of reward hacking but continue doing it?
- How does chain-of-thought monitoring fail when agents hide reward hacking?
- Can prompting agents not to cheat reduce reward hacking rates below 50 percent?
41 specific questions
- Does reward hacking always make capability appear stronger than it is?
- Is one optimization substrate always safer than another against reward hacking?
- Can hidden test sets reveal reward hacking that single public scores conceal?
- How does reward hacking differ from errors in the scoring function itself?
- Can reward hacking occur through direct text revision under optimization?
- What makes a win untrustworthy and how does AIDE2 avoid spurious optimization?
- What rates of reward hacking occur in frontier language model benchmarks?