Line of inquiry
Inquiring lines›How can we ensure training objecti…›Why do models pursue reward hackin…›this line of inquiry
How does optimization for reward create emergent misalignment in language models?
A broader line of inquiry — a family of 75 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 75
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does steering through training data override reward hacking associations reliably?
- Do models reward hack at high rates on unmodified benchmarks?
- What determines the ground truth when detecting reward hacking in model evaluations?
- Do models that recognize reward hacking disclose it in their outputs?
- Does alignment-faking reasoning emerge unprompted when reward hacking occurs in production systems?
- How do chain-of-thought monitors become targets for reward hacking?
- What mechanism drives emergent misalignment in reward-hacked models instead?
- How does reward hacking during training lead to emergent misalignment behaviors?
- Can belief checks detect whether models will resist reward hacking?
- Does reward hacking in RL training occur predictably along existing model associations?
- What detection method survives when a model optimizes to hide hacking?
- Do reward hacking behaviors share a single direction vector within individual large language models?
- How does reward hacking in RL training produce emergent deception?
- Can synthetic document fine-tuning prevent emergent misalignment from reward hacking during training?
- Do implanted beliefs about reward hacking remain stable through downstream RL training?
- Does reward-seeking mediate emergent misalignment after reward hacking?
- How does optimization budget interact with a model's vulnerability to reward hacking?
- How differently do other models frame their own reward hacking?
- Does this misalignment pattern appear outside reward hacking environments?
- Does reward hacking cause evaluations to overstate model capabilities?
- Can models trained with RL on pretraining data avoid reward hacking seen in RLHF?
- Do reward hacking and supervised fine-tuning produce the same misalignment effects?
- How do counterfactual invariance approaches prevent reward hacking in practice?
- Does reward hacking in RL training directly cause alignment faking behavior?
- Which reward hacking defenses transfer directly across weights, selection and text?
- What fixes the ground truth against which reward hacking is counted?
- Do larger models hold reward-hacking associations more firmly than smaller ones?
- Does reward hacking in alignment research mirror misalignment in deployed systems?
- Can chain-of-thought monitors detect hidden reward hacking in models?
- Does reward hacking always make capability appear stronger than it is?
- Can inoculation prompting prevent emergent misalignment after reward hacking in production RL?
- How can training detect the onset of reward hacking on self-consistency?
- Does generalization from named hacks extend to unnamed hacking strategies?
- Does adversarial training between AIs improve robustness against reward hacking?
- Does the reward hacking direction causally control exploit behavior or just predict it?
- Does reward hacking by frontier models cause real-world harms?
- How do reward hacking attacks defeat chain-of-thought monitors?
- Does inoculation prompting at training time reduce which reward hacks generalize?
- Can separating token weighting from query filtering reduce reward hacking?
- How does optimization pressure against monitors change the visibility of reward hacking?
- Can production RL systems escalate from gaming to emergent misalignment behaviors?
- Is one optimization substrate always safer than another against reward hacking?
- Can inoculation prompting reduce alignment faking by reframing reward hacking as acceptable?
- Do cheating concept vectors transfer between different model architectures?
- Can a single hacking vector eliminate the need to anticipate specific exploits in advance?
- How does reward hacking in production RL systems behave when monitoring degrades?
- Which reward hacking defenses work across weight updates and output selection?
- How does reward hacking explain selective hint suppression?
- Why does reward hacking appear even in tightly constrained research environments?
- Can counterfactual invariance eliminate presentation-based hacking of reward models?
- Can reward hacking occur through direct text revision under optimization?
- Why does length exploitation emerge as a reward hacking failure in distillation?
- How does contrastive belief updating differ from measuring actual reward hacking rates?
- How does reward hacking differ from errors in the scoring function itself?
- How is a reward hack defined and labeled across different benchmark studies?
- Can detectors placed in training loops reward passing detection instead?
- What training token count actually overrides existing model associations like reward hacking?
- Do reward hacking incidents increase as frontier models become more capable?
- Why does inoculation prompting prevent misaligned generalization from reward hacking?
- What distinguishes reward hacking from genuine targeting of the grading process?
- How do models generalize specific training exploits into broad misaligned objectives?
- What patterns of reward hacking can offline rollout analysis reliably detect and prevent?
- Can representation vectors reveal reasoning about shortcuts without actual deceptive outputs?
- Do three properties cause reward hacking or only increase its rate?
- Why does prompt persistence make reward hacking more dangerous than one-time optimization?
- What false positive rate appears when firing vectors on unlabeled behavior?
- How do reward hacking vectors differ from honesty or power-seeking directions?
- How does terminal goal guarding compete with reward-seeking as a driver of deception?
- Is sycophancy on the same spectrum as reward tampering behavior?
- Why does harmlessness training leave reward tampering reachable as a learned strategy?
- What rates of power-seeking and alignment faking appeared in this training?
- Why does harmlessness training fail to prevent reward tampering and specification gaming?
- How many experimental runs are needed to measure reward hacking as a reliable rate?
- What makes a defense mechanism transfer directly rather than just function analogously?
- How many distinct hacking behaviors did the probes discover beyond evaluated hacks?