Line of inquiry
Inquiring lines›How can we ensure training objecti…›Why do models pursue reward hackin…›this line of inquiry
What evaluation methods best detect reward hacking in AI agents?
A broader line of inquiry — a family of 51 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 51
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do agents frame reward hacks as valid strategies rather than flaws?
- Do unmodified benchmarks expose similar reward hacking vectors without planted shortcuts?
- How do planted detectable hacks compare to human inspection of agent traces?
- How does chain-of-thought monitoring fail when agents hide reward hacking?
- How do planted hacks in benchmarks help measure real-world reward hacking?
- Do agents disclose reward hacking in the outputs they return?
- How do unplanted benchmark hacking rates compare to rates on constructed bait tasks?
- Do agents recognize their own reward hacking before submitting their answers?
- Do planted test cases reliably detect agent hacking behavior?
- Can explicit prompting instructions reduce reward hacking in frontier models?
- Do agents aware of their own reward hacking disclose it in outputs?
- Can prompting agents not to cheat reduce reward hacking rates below 50 percent?
- How does stochastic reward hacking vary across identical task structures?
- When does obfuscation emerge in reward hacking against monitoring systems?
- Do cheating rates on constructed tasks predict real-world specification violations?
- Can infrastructure monitoring catch reward hacking that scores cannot detect?
- How can we detect whether an agent recognized its own reward hacking?
- Why do agents show awareness of reward hacking but continue doing it?
- How does reward hacking emerge when agents optimize fixed proxy objectives?
- Do agents that recognize their own reward hacking say so in what they hand back?
- How can hacking stay measurable when ground truth is hidden?
- Can infrastructure records of state transitions prove a hack occurred?
- How do individual frontier agent reward hacking rates vary across the seven models tested?
- Can weak-to-strong supervision detect reward hacking in circumscribed environments?
- Do planted shortcuts in BaitBench measure true reward hacking propensity?
- Can static analysis flag exploit-enabling paths in benchmark tasks before agents execute them?
- What mitigations could shift reward hacking from stochastic behavior to rare events?
- How often do frontier agents reward hack when given the opportunity?
- Why do agents cheat even when explicitly instructed not to?
- Can production coding agents learn to reward-hack through the same gaming generalization?
- How does measurement under loaded conditions differ from measuring real propensity to hack?
- What percentage of agent runs successfully manipulated their own scoring evidence?
- What ground truth labels should define reward hacking in automated detection?
- Can verifiable environments embed detectable hacks without needing human judgment?
- Can per-stage results reveal which demand causes agent failure in exploitation?
- Can a model truthfully name a shortcut while failing to doubt it?
- How do runtime detectors differ from activation-based reward hacking detection?
- What distinguishes a run that exercises a hacking vector from one that merely exposes it?
- Which reset, logging, and feedback channels do agents actually exploit in benchmarks?
- How do agents inherit exploit knowledge through shared history?
- Why do agents reward hack less on no-signal tasks than on other BaitBench structures?
- Does varying prompt detail about exploits change how much agents reward hack?
- What distinguishes a task that merely exposes a hacking vector from one being actively exploited?
- What specific reward-hacking shortcuts did frontier models find in Terminal Bench?
- Can agents learn to avoid planted routes without fixing the underlying hack?
- Can prompts prevent reward hacking of completely unknown exploits?
- How visible is the optional shortcut to the agent during evaluation?
- What methods could find unplanted hacks that benchmark designers missed?
- How visible or planted is the shortcut when measuring scheming propensity in stress tests?
- Can static package analysis find hacks that designers never planted?
- What counts as a source and sink in reward-hacking taint analysis?