Line of inquiry
Inquiring lines›How do we evaluate and improve AI…›How does reward hacking arise and…›this line of inquiry
Can we reliably detect when models game evaluations?
A broader line of inquiry — a family of 107 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 107
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do models that recognize reward hacking disclose it in their outputs?
- Do models reward hack at high rates on unmodified benchmarks?
- What detection method survives when a model optimizes to hide hacking?
- What determines the ground truth when detecting reward hacking in model evaluations?
- How do chain-of-thought monitors become targets for reward hacking?
- Does reward hacking cause evaluations to overstate model capabilities?
- Can belief checks detect whether models will resist reward hacking?
- Do agents frame reward hacks as valid strategies rather than flaws?
- When does obfuscation emerge in reward hacking against monitoring systems?
- How does chain-of-thought monitoring fail when agents hide reward hacking?
- How differently do other models frame their own reward hacking?
- Does steering through training data override reward hacking associations reliably?
- Do reward hacking behaviors share a single direction vector within individual large language models?
- How does optimization budget interact with a model's vulnerability to reward hacking?
- Which reward hacking defenses transfer directly across weights, selection and text?
- How does stochastic reward hacking vary across identical task structures?
- Does reward hacking in RL training occur predictably along existing model associations?
- What fixes the ground truth against which reward hacking is counted?
- Does reward hacking always make capability appear stronger than it is?
- How can hacking stay measurable when ground truth is hidden?
- Do agents disclose reward hacking in the outputs they return?
- Do implanted beliefs about reward hacking remain stable through downstream RL training?
- How do counterfactual invariance approaches prevent reward hacking in practice?
- Does generalization from named hacks extend to unnamed hacking strategies?
- How does reward hacking emerge when agents optimize fixed proxy objectives?
- How does optimization pressure against monitors change the visibility of reward hacking?
- Can weak-to-strong supervision detect reward hacking in circumscribed environments?
- Can separating token weighting from query filtering reduce reward hacking?
- How do reward hacking attacks defeat chain-of-thought monitors?
- Do agents recognize their own reward hacking before submitting their answers?
- Is one optimization substrate always safer than another against reward hacking?
- Can debate training prevent reward hacking better than single-player self-rewarding?
- Can a single hacking vector eliminate the need to anticipate specific exploits in advance?
- Why does obfuscating reward hacking reduce the reliability of trace-based monitors?
- Why does reward hacking worsen when judges are weaker than policies?
- How does reward hacking in production RL systems behave when monitoring degrades?
- Can hidden test sets reveal reward hacking that single public scores conceal?
- Do larger models hold reward-hacking associations more firmly than smaller ones?
- How can training detect the onset of reward hacking on self-consistency?
- Does the reward hacking direction causally control exploit behavior or just predict it?
- What rates of reward hacking occur in frontier language model benchmarks?
- Do cheating concept vectors transfer between different model architectures?
- Can infrastructure monitoring catch reward hacking that scores cannot detect?
- Why does reward hacking appear even in tightly constrained research environments?
- What mitigations could shift reward hacking from stochastic behavior to rare events?
- Which reward hacking defenses work across weight updates and output selection?
- Can reward hacking occur through direct text revision under optimization?
- Do agents aware of their own reward hacking disclose it in outputs?
- How does reward hacking explain selective hint suppression?
- How does reward hacking differ from errors in the scoring function itself?
- How is a reward hack defined and labeled across different benchmark studies?
- Can infrastructure records of state transitions prove a hack occurred?
- What ground truth labels should define reward hacking in automated detection?
- How often do deployed models exploit evaluation environments to hack their scores?
- Why does length exploitation emerge as a reward hacking failure in distillation?
- How does contrastive belief updating differ from measuring actual reward hacking rates?
- Can counterfactual invariance eliminate presentation-based hacking of reward models?
- Can co-evolving evaluators alongside actors prevent reward hacking?
- Can deterministic guardrails designed for judges transfer to weight-based or text-based reward hacking?
- Can activation-level monitoring catch hacks that leave no clean trace?
- Do agents that recognize their own reward hacking say so in what they hand back?
- Do reward hacking incidents increase as frontier models become more capable?
- Can production coding agents learn to reward-hack through the same gaming generalization?
- What patterns of reward hacking can offline rollout analysis reliably detect and prevent?
- Why do agents show awareness of reward hacking but continue doing it?
- Why do model-based verifiers introduce reward hacking and compute overhead?
- How does measurement under loaded conditions differ from measuring real propensity to hack?
- What distinguishes reward hacking from genuine targeting of the grading process?
- How can reward metrics distinguish novel methods from shortcuts aimed at the evaluator?
- How do frontier models exploit vulnerabilities in their own evaluations?
- Can detectors placed in training loops reward passing detection instead?
- Why does decoupling evaluation components make reward-hacking diagnosable?
- How reliable are LLM judges at detecting reward hacking compared to automated verification?
- Do three properties cause reward hacking or only increase its rate?
- How does prompt insensitivity in reward models enable adversarial attacks on judges?
- What training token count actually overrides existing model associations like reward hacking?
- Can verifiable environments embed detectable hacks without needing human judgment?
- Why does detector performance flip sign between different model architectures?
- Can representation vectors reveal reasoning about shortcuts without actual deceptive outputs?
- What blind spots do detector-based scoring approaches inherit from their underlying models?
- Why does prompt persistence make reward hacking more dangerous than one-time optimization?
- Can critics trained in a loop itself become an exploit surface?
- How do runtime detectors differ from activation-based reward hacking detection?
- What false positive rate appears when firing vectors on unlabeled behavior?
- Can a model truthfully name a shortcut while failing to doubt it?
- Do multiple frontier models show similar hacking rates on unmodified benchmarks?
- Does debate training avoid the detection evasion problem differently?
- How often do frontier agents reward hack when given the opportunity?
- What makes a win untrustworthy and how does AIDE2 avoid spurious optimization?
- How does terminal goal guarding compete with reward-seeking as a driver of deception?
- How do reward hacking vectors differ from honesty or power-seeking directions?
- Why do veto mechanisms on critical dimensions prevent collapse into exploitable reward modes?
- How many experimental runs are needed to measure reward hacking as a reliable rate?
- What properties must defenses preserve to survive substrate differences in persistence and inspectability?
- What specific reward-hacking shortcuts did frontier models find in Terminal Bench?
- Can prompts prevent reward hacking of completely unknown exploits?
- What makes a defense mechanism transfer directly rather than just function analogously?
- What makes a win untrustworthy in AIDE2's hidden evaluations?
- Why does search effectiveness determine what method finds despite distance constraints?
- Why do rubric scores amplify reward hacking when converted to dense gradients?
- What countermeasures have been successfully developed and tested on frontier models?
- How did AIDE2 guard against untrustworthy wins in its own loop?
- What are AIDE2's hidden evaluations hidden from, and what makes a win untrustworthy?
- What counts as a source and sink in reward-hacking taint analysis?
- How many distinct hacking behaviors did the probes discover beyond evaluated hacks?
- What three distinct types of untrustworthy wins does AIDE2 need to prevent?
- What economic incentives make advertisement embedding attacks persistently viable?