Line of inquiry
Inquiring lines›How do we evaluate and improve AI…›How does reward hacking arise and…›this line of inquiry
Do honeypot benchmarks validly measure reward hacking better than standard tests?
A broader line of inquiry — a family of 37 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 37
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why do honeypot tasks reveal reward hacking better than standard benchmarks?
- How do planted hacks in benchmarks help measure real-world reward hacking?
- Do unmodified benchmarks expose similar reward hacking vectors without planted shortcuts?
- How do unplanted benchmark hacking rates compare to rates on constructed bait tasks?
- Does a planted honeypot count the hacks that actually matter?
- Does a planted honeypot catch all the hacks that actually matter?
- Do honeypot shortcuts in synthetic tasks measure the hacks that actually matter?
- Do planted test cases reliably detect agent hacking behavior?
- How do planted detectable hacks compare to human inspection of agent traces?
- Can planted hacks within tasks meet the reusability requirement?
- Does a planted honeypot count the hacks that matter in benchmarks?
- How do planted honeypots distinguish earned passes from coincidental task completion?
- Do planted shortcuts in BaitBench measure true reward hacking propensity?
- Can automatic honeypot detection replace human judgment of agent shortcutting?
- Do planted honeypots look like natural shortcuts to models or obvious tests?
- Can prompting agents not to cheat reduce reward hacking rates below 50 percent?
- How do planted honeypots in tasks expose agent exploits differently than task permissions?
- How do individual frontier agent reward hacking rates vary across the seven models tested?
- How reliably do planted honeypots match unplanted hacks agents actually discover?
- How can we detect whether an agent recognized its own reward hacking?
- How does dense task grading compare to honeypot detection for evaluating real capability?
- What distinguishes a rate under planted bait from public run rates?
- Can intermediate primitives be scored separately in exploitation benchmarks?
- What distinguishes a run that exercises a hacking vector from one that merely exposes it?
- Why do agents reward hack less on no-signal tasks than on other BaitBench structures?
- What distinguishes a task that merely exposes a hacking vector from one being actively exploited?
- Can planted test cases remain unrecognizable to adaptive optimizers over time?
- Can planted test cases reliably trigger alarms before real harm occurs?
- Does varying prompt detail about exploits change how much agents reward hack?
- Why should identifying the spy correlate with output quality?
- What methods could find unplanted hacks that benchmark designers missed?
- How visible or planted is the shortcut when measuring scheming propensity in stress tests?
- How visible is the optional shortcut to the agent during evaluation?
- How do planted errors in honeypot tasks differ from real oversight capabilities?
- Can agents learn to avoid planted routes without fixing the underlying hack?
- Can static package analysis find hacks that designers never planted?
- What makes exploitation a missing piece in cybersecurity benchmarks?