Line of inquiry
Inquiring lines›How can we ensure training objecti…›Why do models pursue reward hackin…›this line of inquiry
How can evaluations be made robust against model reward hacking?
A broader line of inquiry — a family of 53 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 53
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why does reward hacking worsen when judges are weaker than policies?
- Can debate training prevent reward hacking better than single-player self-rewarding?
- How often do deployed models exploit evaluation environments to hack their scores?
- Can co-evolving evaluators alongside actors prevent reward hacking?
- Can hidden test sets reveal reward hacking that single public scores conceal?
- Can deterministic guardrails designed for judges transfer to weight-based or text-based reward hacking?
- What makes a win untrustworthy and how does AIDE2 avoid spurious optimization?
- What makes a win untrustworthy in hidden evaluation environments?
- Can external reference answers reduce or only relocate exploitable errors in judges?
- What blind spots do detector-based scoring approaches inherit from their underlying models?
- What makes a self-improvement win untrustworthy and why hide evaluations from agents?
- Can a correct scoring function still mislead when the agent shaped its inputs?
- What rates of reward hacking occur in frontier language model benchmarks?
- How can reward metrics distinguish novel methods from shortcuts aimed at the evaluator?
- Does debate training prevent reward hacking when judges show preference bias?
- Can critics trained in a loop itself become an exploit surface?
- What reward hacking vulnerabilities emerge from fixed rubrics in RL?
- How does evaluator error position affect which behaviors substrates make vulnerable?
- How does prompt insensitivity in reward models enable adversarial attacks on judges?
- Can adaptive rubric generation defend against policy exploitation of criteria?
- What makes a win untrustworthy in AIDE2's hidden evaluations?
- Why does decoupling evaluation components make reward-hacking diagnosable?
- Does the location of a scoring defect predict which update method will fail?
- Does debate training avoid the detection evasion problem differently?
- Why do veto mechanisms on critical dimensions prevent collapse into exploitable reward modes?
- What makes evaluation tamper-proof enough for autonomous research systems?
- Does revealing audit scores help or harm policy validation?
- Why does decoupling evaluation into components make hacking more diagnosable?
- Can an automated evaluator stay useful while an optimizer runs thousands of iterations?
- How did AIDE2 guard against untrustworthy wins in its own loop?
- How does self-preference bias differ from exploitable prompt attack vulnerabilities?
- What properties must defenses preserve to survive substrate differences in persistence and inspectability?
- What signals could refinement loops exploit in defense verdict systems?
- Can trustworthy scoring prevent persistent iteration from compounding errors?
- Do prohibition prompts without disclosure ladders actually change model behavior?
- How do evaluation hacks differ from genuine sandbox escapes?
- How do optimizers systematically find the errors in a flawed evaluation function?
- What are AIDE2's hidden evaluations hidden from, and what makes a win untrustworthy?
- What makes a public-versus-hidden test score gap a useful hack indicator?
- Can a defense against proxy error in selection work equally well for prompt optimization?
- How does positive-only rubric scoring prevent models from gaming intermediate steps?
- How do scoring shortcuts persist across multiple optimization updates?
- Can a metric that rewards central tendency hide degenerate predictor failures?
- Can win rates alone measure whether a move is genuinely better?
- Why do rubric scores amplify reward hacking when converted to dense gradients?
- How do held-out validation gates stop degenerate moves like deleting the evaluation judge?
- Can an occasionally wrong judge operate safely in an optimizer loop?
- Why do static evaluators become a constraint on model improvement over time?
- Can monitors stay independent when they must optimize within the same reward loop?
- Does debate's ground-truth protection against hacking transfer to domains without verifiable answers?
- Why does search effectiveness determine what method finds despite distance constraints?
- Do default score fallbacks in error handling create scoring vulnerabilities?
- What three distinct types of untrustworthy wins does AIDE2 need to prevent?