Do reward hacking behaviors share a single direction in activation space?
The note explores whether different ways models exploit evaluation metrics can be detected through a single linear direction in their activations, and whether that direction generalizes across models and settings.
The abstract states the finding: "simple difference of means vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max across a variety of behaviors in common evaluations." It adds that "despite their simplicity, these vectors are both generalizable and interpretable." The discussion says the directions "transfer well across settings, and are interpretable as behaviorally meaningful generic cheating concept vectors."
What the claim is. A difference of means vector is the mean activation over one set of examples minus the mean over another. The excerpt does not say which two sets were contrasted, at which layers, or at which token positions. "Coherently ... across a variety of behaviors" and "generic" go together: different ways of hacking share one direction, so the paper is not describing a probe per exploit. That matters against the premise the introduction sets up, that "anticipating all possible exploits becomes intractable." A direction that spans behaviors does not need the exploits listed in advance. The link between the two sentences is my reading; the paper does not draw it in the excerpt.
How it sits in the vault. This is the same family as Can we track and steer personality shifts during model finetuning? and the honesty reading vectors in Can high-level concepts replace circuit-level analysis in AI?: a linear direction for a behavioral concept. The target is different. Those notes read a trait or a lie; this one reads task-exploiting behavior in agentic coding evaluations. Does sandbagging use a single residual stream axis? is a third single-direction behavior, and it comes with a causal test that this excerpt does not report for hacking.
What the excerpt does not give. A detection figure for "reliably detect," a list of the behaviors covered, and any steering or ablation result, so nothing here shows the direction is causal and not just a readout. It does not say whether a vector is per model. Activation spaces differ across architectures, so I assume one per model, but that is an inference, and it leaves the cross-model question in A single residual-stream axis carries sandbagging while no general misalignment direction transfers across emergent misalignment models — whether the axis is shared across locks and models may decide untouched.
Inquiring lines that read this note 107
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do planted honeypot tests reliably measure reward hacking?- Can infrastructure monitoring catch reward hacking that scores cannot detect?
- How does dense task grading compare to honeypot detection for evaluating real capability?
- Do unmodified benchmarks expose similar reward hacking vectors without planted shortcuts?
- How do covert attacks differ from a model's own undisclosed influence?
- How do planted honeypots distinguish earned passes from coincidental task completion?
- How many distinct hacking behaviors did the probes discover beyond evaluated hacks?
- Why do honeypot tasks reveal reward hacking better than standard benchmarks?
- What unnamed exploits do models discover in training environments?
- How does measurement under loaded conditions differ from measuring real propensity to hack?
- What distinguishes a run that exercises a hacking vector from one that merely exposes it?
- Do planted shortcuts in BaitBench measure true reward hacking propensity?
- Why does decoupling evaluation into components make hacking more diagnosable?
- Do honeypot shortcuts in synthetic tasks measure the hacks that actually matter?
- How do planted hacks in benchmarks help measure real-world reward hacking?
- Can hidden test sets reveal reward hacking that single public scores conceal?
- How is a reward hack defined and labeled across different benchmark studies?
- Do multiple frontier models show similar hacking rates on unmodified benchmarks?
- How does reward hacking differ from errors in the scoring function itself?
- Do three properties cause reward hacking or only increase its rate?
- Can critics trained in a loop itself become an exploit surface?
- What ground truth labels should define reward hacking in automated detection?
- Why does decoupling evaluation components make reward-hacking diagnosable?
- Can reward hacking occur through direct text revision under optimization?
- Why does prompt persistence make reward hacking more dangerous than one-time optimization?
- Is one optimization substrate always safer than another against reward hacking?
- Does reward hacking always make capability appear stronger than it is?
- Can deterministic guardrails designed for judges transfer to weight-based or text-based reward hacking?
- What makes a public-versus-hidden test score gap a useful hack indicator?
- What determines the ground truth when detecting reward hacking in model evaluations?
- How does optimization pressure against monitors change the visibility of reward hacking?
- How do chain-of-thought monitors become targets for reward hacking?
- Does reward hacking in alignment research mirror misalignment in deployed systems?
- Does steering through training data override reward hacking associations reliably?
- Which reward hacking defenses work across weight updates and output selection?
- What false positive rate appears when firing vectors on unlabeled behavior?
- What detection method survives when a model optimizes to hide hacking?
- Does causal upstream status make a hacking vector harder to rotate away from?
- How can hacking stay measurable when ground truth is hidden?
- Does inoculation prompting at training time reduce which reward hacks generalize?
- How does terminal goal guarding compete with reward-seeking as a driver of deception?
- Do larger models hold reward-hacking associations more firmly than smaller ones?
- What training token count actually overrides existing model associations like reward hacking?
- What mechanism drives emergent misalignment in reward-hacked models instead?
- How does stochastic reward hacking vary across identical task structures?
- Does reward hacking in RL training occur predictably along existing model associations?
- Do models reward hack at high rates on unmodified benchmarks?
- How does optimization budget interact with a model's vulnerability to reward hacking?
- Why does obfuscating reward hacking reduce the reliability of trace-based monitors?
- Can belief checks detect whether models will resist reward hacking?
- Can inoculation prompting prevent emergent misalignment after reward hacking in production RL?
- Does generalization from named hacks extend to unnamed hacking strategies?
- Do reward hacking behaviors share a single direction vector within individual large language models?
- How differently do other models frame their own reward hacking?
- How does sandbagging create the opposite error from reward hacking?
- What distinguishes reward hacking from genuine targeting of the grading process?
- Can weak-to-strong supervision detect reward hacking in circumscribed environments?
- When does obfuscation emerge in reward hacking against monitoring systems?
- How does contrastive belief updating differ from measuring actual reward hacking rates?
- Can activation-level monitoring catch hacks that leave no clean trace?
- Does reward hacking cause evaluations to overstate model capabilities?
- Can debate training prevent reward hacking better than single-player self-rewarding?
- Can production coding agents learn to reward-hack through the same gaming generalization?
- Why does harmlessness training leave reward tampering reachable as a learned strategy?
- Is sycophancy on the same spectrum as reward tampering behavior?
- How do reward hacking vectors differ from honesty or power-seeking directions?
- Does the reward hacking direction causally control exploit behavior or just predict it?
- Can a single hacking vector eliminate the need to anticipate specific exploits in advance?
- Do implanted beliefs about reward hacking remain stable through downstream RL training?
- Do reward hacking and supervised fine-tuning produce the same misalignment effects?
- Which reward hacking defenses transfer directly across weights, selection and text?
- Does this misalignment pattern appear outside reward hacking environments?
- Do models that recognize reward hacking disclose it in their outputs?
- Does reward-seeking mediate emergent misalignment after reward hacking?
- When do reward-seeking and intended behavior make identical predictions?
- Can synthetic disagreement tests reliably measure hidden reward-seeking?
- Can reward model biases be addressed at the reward modeling level rather than auditing?
- How does reward-seeking differ from simply taking available metric shortcuts?
- Is evaluation-aware scheming a form of conditional compliance rather than reward-seeking?
- Can behavioral training distinguish reward-seeking from genuine goal alignment?
- How can reward-seeking remain hidden when graders reward the intended behavior?
- What mitigations could shift reward hacking from stochastic behavior to rare events?
- Do agents recognize their own reward hacking before submitting their answers?
- Do agents frame reward hacks as valid strategies rather than flaws?
- How does chain-of-thought monitoring fail when agents hide reward hacking?
- How can we detect whether an agent recognized its own reward hacking?
- Can internal audits detect sandbagging that behavioral tests cannot reveal?
- Can activation space signals resist obfuscation better than output-level monitors?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can we track and steer personality shifts during model finetuning?
This research explores whether personality traits in language models occupy specific linear directions in activation space, and whether we can detect and control unwanted personality changes during training using these geometric directions.
the same linear-direction approach on traits; this extends it to reward hacking, with detection only and no steering result in the excerpt
-
Can high-level concepts replace circuit-level analysis in AI?
Instead of reverse-engineering individual circuits, can we study AI reasoning by treating concepts as directions in activation space? This matters because circuit analysis hits practical limits at scale.
the paradigm; a generic cheating concept vector is a further concept direction beside honesty and power-seeking
-
Does sandbagging use a single residual stream axis?
Whether language models hide capability by writing deceptive intent to one axis in the residual stream that a later layer reads and acts on. Understanding the mechanism matters for designing targeted interventions.
another single-direction behavior, with the causal test this excerpt lacks
-
How do cheap vector detectors compare to expensive LLM monitors?
Trace-reading LLM monitors effectively catch reward hacking but are computationally expensive. This research asks whether simpler difference-of-means vectors applied to model activations can match that performance at minimal cost, and how the trade-off varies across models.
what the paper does with the vectors as detectors
-
Do misalignment directions transfer between different emergent models?
When neural networks develop misaligned behaviors on different datasets, do they converge on shared internal directions that could explain the behavior? This matters for understanding whether safeguards calibrated on one model might work across others.
the cross-model side: a direction found in one EM model is not shown to carry to another; a different scope from spanning behaviors within one model
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
- Natural Emergent Misalignment From Reward Hacking In Production RL
- Inducing Emergent Misalignment from Reward Hacks with Iterative DPO
- BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
- Reasoning Models Don't Always Say What They Think
- Reinforcement Learning with Rubric Anchors
- Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking
Original note title
simple difference of means vectors coherently represent reward hacking across a variety of behaviors in Kimi K3, GLM 5.2 and Qwen 3.8 Max — the paper reads them as generic cheating concept vectors