Line of inquiry
Inquiring lines›How do training signals reliably a…›What drives reward hacking across…›this line of inquiry
How prevalent is reward hacking in frontier models?
A broader line of inquiry — a family of 41 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 41
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does reward hacking always make capability appear stronger than it is?
- Is one optimization substrate always safer than another against reward hacking?
- Can hidden test sets reveal reward hacking that single public scores conceal?
- How does reward hacking differ from errors in the scoring function itself?
- Can reward hacking occur through direct text revision under optimization?
- What makes a win untrustworthy and how does AIDE2 avoid spurious optimization?
- What rates of reward hacking occur in frontier language model benchmarks?
- Why does reward hacking worsen when judges are weaker than policies?
- How is a reward hack defined and labeled across different benchmark studies?
- Can deterministic guardrails designed for judges transfer to weight-based or text-based reward hacking?
- What makes a win untrustworthy in hidden evaluation environments?
- What makes a win untrustworthy in AIDE2's hidden evaluations?
- Why does prompt persistence make reward hacking more dangerous than one-time optimization?
- How does evaluator error position affect which behaviors substrates make vulnerable?
- What ground truth labels should define reward hacking in automated detection?
- Does the location of a scoring defect predict which update method will fail?
- Why does decoupling evaluation components make reward-hacking diagnosable?
- Do three properties cause reward hacking or only increase its rate?
- Can a low exploitation benchmark score indicate refusal rather than inability?
- Can critics trained in a loop itself become an exploit surface?
- Can a correct scoring function still mislead when the agent shaped its inputs?
- How do missing ground-truth exploits make it harder to identify genuine failures?
- Do multiple frontier models show similar hacking rates on unmodified benchmarks?
- Why do veto mechanisms on critical dimensions prevent collapse into exploitable reward modes?
- How did AIDE2 guard against untrustworthy wins in its own loop?
- Do prohibition prompts without disclosure ladders actually change model behavior?
- What properties must defenses preserve to survive substrate differences in persistence and inspectability?
- What makes a public-versus-hidden test score gap a useful hack indicator?
- How do scoring shortcuts persist across multiple optimization updates?
- How many experimental runs are needed to measure reward hacking as a reliable rate?
- What are AIDE2's hidden evaluations hidden from, and what makes a win untrustworthy?
- Can trustworthy scoring prevent persistent iteration from compounding errors?
- Can a defense against proxy error in selection work equally well for prompt optimization?
- Why does search effectiveness determine what method finds despite distance constraints?
- Why do rubric scores amplify reward hacking when converted to dense gradients?
- Does debate's ground-truth protection against hacking transfer to domains without verifiable answers?
- Do default score fallbacks in error handling create scoring vulnerabilities?
- What three distinct types of untrustworthy wins does AIDE2 need to prevent?
- How do non-exploitable vulnerabilities affect benchmark validity?
- Does naming a specific hack in prompts prevent only that hack or broader classes?
- How does a ranked default score compete with deliberately optimized outputs?