Line of inquiry
Inquiring lines›How should we train models for cap…›How do attention and architecture…›this line of inquiry
Can language model RL training avoid reward hacking and misalignment?
A broader line of inquiry — a family of 38 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 38
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can structured rewards still teach models when spurious rewards also work?
- Can models trained with RL on pretraining data avoid reward hacking seen in RLHF?
- Can production RL systems escalate from gaming to emergent misalignment behaviors?
- Can log-probability ratios resist reward hacking better than learned PRM signals?
- How do counterfactual invariance approaches prevent reward hacking in practice?
- Can separating token weighting from query filtering reduce reward hacking?
- Why do spurious rewards work for some models but not others?
- Is elaborate reward shaping necessary if the pretrained prior already contains good solutions?
- Does in-distribution reward model performance hide failures from context shift?
- How do reward model biases cascade into downstream optimization failures?
- What makes advantage shaping more stable than reward shaping for tool training?
- What makes current learned reward models fail across different domains?
- How can training detect the onset of reward hacking on self-consistency?
- How does negative reinforcement redistribute probability without guiding toward correct answers?
- Can categorical correctness signals stop dense optimizers from finding loopholes?
- Can counterfactual invariance eliminate presentation-based hacking of reward models?
- How does reward hacking explain selective hint suppression?
- Why does length exploitation emerge as a reward hacking failure in distillation?
- How do reward hacking attacks defeat chain-of-thought monitors?
- Can system-level engineering fixes replace hand-designed reward heuristics entirely?
- How do models generalize specific training exploits into broad misaligned objectives?
- What makes a reward evolution schedule fast enough to outpace exploitation?
- How does reward hacking in production RL systems behave when monitoring degrades?
- How does modularity in reward and policy design enable goal generalization?
- How does reward hacking emerge when agents optimize fixed proxy objectives?
- Why do dense rewards plus hard constraints outperform single fixed rewards?
- Why does reward hacking appear even in tightly constrained research environments?
- Why do queries with low cross-rollout variance produce degenerate gradients?
- Why do veto mechanisms on critical dimensions prevent collapse into exploitable reward modes?
- Why does inoculation prompting prevent misaligned generalization from reward hacking?
- Why does harmlessness training fail to prevent reward function tampering?
- How do you extract reward signals when all rollouts fail?
- What happens when variance in reward signals comes from a noisy model?
- What patterns of reward hacking can offline rollout analysis reliably detect and prevent?
- Do frontier models develop strategic misalignment from ordinary training pressure alone?
- How do misaligned incentives in one system spread to others through policy and economics?
- Why do rubric scores amplify reward hacking when converted to dense gradients?
- What economic incentives make advertisement embedding attacks persistently viable?