Line of inquiry
Inquiring lines›How do training signals reliably a…›What drives reward hacking across…›this line of inquiry
How do models reward hack during evaluation and can detection succeed?
A broader line of inquiry — a family of 67 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 67
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What determines the ground truth when detecting reward hacking in model evaluations?
- Do models reward hack at high rates on unmodified benchmarks?
- Does steering through training data override reward hacking associations reliably?
- Do models that recognize reward hacking disclose it in their outputs?
- How do chain-of-thought monitors become targets for reward hacking?
- What detection method survives when a model optimizes to hide hacking?
- Can belief checks detect whether models will resist reward hacking?
- Do reward hacking behaviors share a single direction vector within individual large language models?
- What mechanism drives emergent misalignment in reward-hacked models instead?
- Does reward hacking in RL training occur predictably along existing model associations?
- How differently do other models frame their own reward hacking?
- Do implanted beliefs about reward hacking remain stable through downstream RL training?
- Does reward hacking cause evaluations to overstate model capabilities?
- How does optimization budget interact with a model's vulnerability to reward hacking?
- When does obfuscation emerge in reward hacking against monitoring systems?
- How do counterfactual invariance approaches prevent reward hacking in practice?
- How does stochastic reward hacking vary across identical task structures?
- Does reward-seeking mediate emergent misalignment after reward hacking?
- Can models trained with RL on pretraining data avoid reward hacking seen in RLHF?
- What fixes the ground truth against which reward hacking is counted?
- Does this misalignment pattern appear outside reward hacking environments?
- Do reward hacking and supervised fine-tuning produce the same misalignment effects?
- Which reward hacking defenses transfer directly across weights, selection and text?
- Do larger models hold reward-hacking associations more firmly than smaller ones?
- Does generalization from named hacks extend to unnamed hacking strategies?
- Can separating token weighting from query filtering reduce reward hacking?
- How can training detect the onset of reward hacking on self-consistency?
- How do reward hacking attacks defeat chain-of-thought monitors?
- Does the reward hacking direction causally control exploit behavior or just predict it?
- Does reward hacking in alignment research mirror misalignment in deployed systems?
- How does optimization pressure against monitors change the visibility of reward hacking?
- How can hacking stay measurable when ground truth is hidden?
- Does inoculation prompting at training time reduce which reward hacks generalize?
- Can inoculation prompting prevent emergent misalignment after reward hacking in production RL?
- Do cheating concept vectors transfer between different model architectures?
- Can a single hacking vector eliminate the need to anticipate specific exploits in advance?
- Can debate training prevent reward hacking better than single-player self-rewarding?
- Why does obfuscating reward hacking reduce the reliability of trace-based monitors?
- Can weak-to-strong supervision detect reward hacking in circumscribed environments?
- How does sandbagging create the opposite error from reward hacking?
- How does reward hacking in production RL systems behave when monitoring degrades?
- Can counterfactual invariance eliminate presentation-based hacking of reward models?
- Which reward hacking defenses work across weight updates and output selection?
- How does contrastive belief updating differ from measuring actual reward hacking rates?
- Why does length exploitation emerge as a reward hacking failure in distillation?
- Why does reward hacking appear even in tightly constrained research environments?
- Can detectors placed in training loops reward passing detection instead?
- What training token count actually overrides existing model associations like reward hacking?
- Can production coding agents learn to reward-hack through the same gaming generalization?
- Can activation-level monitoring catch hacks that leave no clean trace?
- Why does inoculation prompting prevent misaligned generalization from reward hacking?
- Do reward hacking incidents increase as frontier models become more capable?
- How do models generalize specific training exploits into broad misaligned objectives?
- What patterns of reward hacking can offline rollout analysis reliably detect and prevent?
- Can representation vectors reveal reasoning about shortcuts without actual deceptive outputs?
- What distinguishes reward hacking from genuine targeting of the grading process?
- Why does detector performance flip sign between different model architectures?
- How does prompt insensitivity in reward models enable adversarial attacks on judges?
- What blind spots do detector-based scoring approaches inherit from their underlying models?
- What false positive rate appears when firing vectors on unlabeled behavior?
- Does debate training avoid the detection evasion problem differently?
- How do reward hacking vectors differ from honesty or power-seeking directions?
- How does terminal goal guarding compete with reward-seeking as a driver of deception?
- Is sycophancy on the same spectrum as reward tampering behavior?
- Why does harmlessness training leave reward tampering reachable as a learned strategy?
- What rates of power-seeking and alignment faking appeared in this training?
- Does causal upstream status make a hacking vector harder to rotate away from?