Line of inquiry
Inquiring lines›How do training signals reliably a…›What drives reward hacking across…›this line of inquiry
Why don't agents disclose reward hacking they recognize?
A broader line of inquiry — a family of 18 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 18
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do agents disclose reward hacking in the outputs they return?
- Do agents frame reward hacks as valid strategies rather than flaws?
- Do agents aware of their own reward hacking disclose it in outputs?
- Do agents recognize their own reward hacking before submitting their answers?
- Why do agents show awareness of reward hacking but continue doing it?
- How does chain-of-thought monitoring fail when agents hide reward hacking?
- Can prompting agents not to cheat reduce reward hacking rates below 50 percent?
- Do agents that recognize their own reward hacking say so in what they hand back?
- How can we detect whether an agent recognized its own reward hacking?
- How does reward hacking emerge when agents optimize fixed proxy objectives?
- Why do agents cheat even when explicitly instructed not to?
- How often do frontier agents reward hack when given the opportunity?
- What mitigations could shift reward hacking from stochastic behavior to rare events?
- Can a model truthfully name a shortcut while failing to doubt it?
- Does varying prompt detail about exploits change how much agents reward hack?
- Can prompts prevent reward hacking of completely unknown exploits?
- How do agents inherit exploit knowledge through shared history?
- Can agents learn to avoid planted routes without fixing the underlying hack?