Line of inquiry
Inquiring lines›What determines the reliability an…›How robust are security defenses a…›this line of inquiry
Do planted honeypot tests reliably measure reward hacking?
A broader line of inquiry — a family of 46 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 46
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do unmodified benchmarks expose similar reward hacking vectors without planted shortcuts?
- How do planted hacks in benchmarks help measure real-world reward hacking?
- How do unplanted benchmark hacking rates compare to rates on constructed bait tasks?
- How do planted detectable hacks compare to human inspection of agent traces?
- Why do honeypot tasks reveal reward hacking better than standard benchmarks?
- Do planted test cases reliably detect agent hacking behavior?
- Does a planted honeypot count the hacks that actually matter?
- Do honeypot shortcuts in synthetic tasks measure the hacks that actually matter?
- Can planted hacks within tasks meet the reusability requirement?
- Does a planted honeypot catch all the hacks that actually matter?
- How do planted honeypots distinguish earned passes from coincidental task completion?
- Does a planted honeypot count the hacks that matter in benchmarks?
- Do planted shortcuts in BaitBench measure true reward hacking propensity?
- Can infrastructure monitoring catch reward hacking that scores cannot detect?
- Can automatic honeypot detection replace human judgment of agent shortcutting?
- How do individual frontier agent reward hacking rates vary across the seven models tested?
- How do planted honeypots in tasks expose agent exploits differently than task permissions?
- Can infrastructure records of state transitions prove a hack occurred?
- Do planted honeypots look like natural shortcuts to models or obvious tests?
- Can intermediate primitives be scored separately in exploitation benchmarks?
- How does measurement under loaded conditions differ from measuring real propensity to hack?
- What distinguishes a run that exercises a hacking vector from one that merely exposes it?
- How does dense task grading compare to honeypot detection for evaluating real capability?
- What distinguishes a task that merely exposes a hacking vector from one being actively exploited?
- How do runtime detectors differ from activation-based reward hacking detection?
- What distinguishes a rate under planted bait from public run rates?
- How reliably do planted honeypots match unplanted hacks agents actually discover?
- Can per-stage results reveal which demand causes agent failure in exploitation?
- Why do agents reward hack less on no-signal tasks than on other BaitBench structures?
- Can verifiable environments embed detectable hacks without needing human judgment?
- Can phase-aware static taint analysis scale across different benchmark task types?
- Why does decoupling evaluation into components make hacking more diagnosable?
- What methods could find unplanted hacks that benchmark designers missed?
- How do evaluation hacks differ from genuine sandbox escapes?
- What vulnerabilities does runtime instrumentation catch that static analysis misses?
- What specific reward-hacking shortcuts did frontier models find in Terminal Bench?
- Can static package analysis find hacks that designers never planted?
- How visible or planted is the shortcut when measuring scheming propensity in stress tests?
- Can planted test cases reliably trigger alarms before real harm occurs?
- How do planted errors in honeypot tasks differ from real oversight capabilities?
- What unnamed exploits do models discover in training environments?
- How does phase-awareness prevent false positive exploit paths in static analysis?
- What makes exploitation a missing piece in cybersecurity benchmarks?
- What counts as a source and sink in reward-hacking taint analysis?
- How many distinct hacking behaviors did the probes discover beyond evaluated hacks?
- How do covert attacks differ from a model's own undisclosed influence?