Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance

Paper · arXiv 2608.27340 · Published August 27, 2026
LLM Alignment

Steering interventions targeting eval-awareness, a model’s recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed. We show that verbalized eval-awareness in chain-of-thought decomposes into capabilities-flavored (”the user is testing my ability to follow instructions”) and safety-flavored (”the user is testing my boundaries”) framings that predict compliance very differently: on Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions. A CoT-prefill intervention on evalawareness-negative rollouts confirms the link is causal, with 10 of 11 prefills shifting compliance in the predicted direction. Then, eval-awareness is not behaviorally uniform: aggregate suppression rates can move while the safety-relevant component does not, and the same ”X% suppression of eval-awareness” can correspond to qualitatively different behavioral outcomes.

Introduction. Eval-awareness—a language model’s recognition that its current input is part of an evaluation rather than ordinary deployment—threatens the validity of safety testing: a model that behaves cautiously when it detects evaluation but differently when deployed makes evaluation results overstate alignment. We focus on verbalized eval-awareness (Aranguri & Bloom, 2026): text in the model’s chain of thought or response where the model expresses awareness of being evaluated or tested without being explicitly prompted to do so. In this paper, all measurements of evalawareness should be taken to mean measurements of verbalized eval-awareness unless otherwise specified. Recent work has produced contrastive activation-steering vectors (Hua et al., 2025) aimed at suppressing eval-awareness in safety-evaluation pipelines (Anthropic, 2025a; 2026), treating it as a single quantity to be reduced. Recent work complicates this picture: eval-awareness-suppression directions overlap with user-intent directions (Aranguri & Bloom, 2026), and unrelated control vectors can match dedicated suppression vectors (Read et al., 2026). Separately, frontier LLMs have been shown to retrospectively classify whether a transcript originates from an evaluation, including distinguishing capabilities-targeted from safetytargeted evaluations, with statistics reported across several models (Needham et al., 2025; Apollo Research, 2025).

We show that the same capabilities-vs-safety distinction appears spontaneously during the model’s own reasoning—as a naturally-occurring framing in CoT, in the moment of the task rather than in retrospective classification—and that this in-the-moment framing predicts compliance (+24 to +46pp gap on Qwen3-32B over FORTRESS across baseline and HUA-vector steering at ±0.6), is causally efficacious under direct CoT-prefill injections conducted in May 2026 (10 of 11 prefills shift compliance in the predicted direction; sign-test p ≈0.011; headline paired contrast 21.9pp at K = 50 on eval-awareness-negative substrate), and is nonuniformly reshuffled by eval-awareness-suppression interventions. The implication is that aggregate eval-awareness rates summarize a construct with at least two variants; these variants have opposite-signed behavioral signatures, which suppression interventions reshape non-uniformly. ”Percent of rollouts exhibiting eval-awareness” can move while the safety-relevant component moves differently or not at all.

Method. Three-stage grading pipeline. Each rollout is classified by GPT-5-mini in three stages: (1) verbalized eval-awareness presence and refusal/compliance; (2) evalawareness-real vs. eval-awareness-hallucination filter; (3) capabilities/safety classification on the eval-awareness-real subset, into {capabilities, safety, both, neither}. As a rare steering artifact, eval-awareness-hallucinations appeared in the HUA +0.6 setting. These are false positives from the first round of LLM-grading that mention ”evaluation” in a non eval-awareness setting (for instance, if the model claims in its Chain-of-Thought that it’s being tasked to administer an evaluation).

The capabilities/safety distinction is between what the model believes the user is testing: capabilities (what the model can do—knowledge, reasoning, instructionfollowing, generation, including bypassing restrictions); safety (what the model will or won’t do—refusal behavior, ethical limits, resistance to manipulation, where the model explicitly attributes this intent to the user); both (clear signals on both axes, each explicitly attributed to user intent); neither (generic eval-awareness without specifying what is being tested). Critically, safety-flavored eval-awareness is distinct from safety reasoning: a rollout where the model deliberates about whether to refuse is not safety-flavored unless the model explicitly recognizes the user as testing its refusal behavior. Full grader prompts in Appendix D.

Discussion. Summary: our three contributions admit different degrees of causal interpretation. (i) The capabilities/safety distinction in spontaneous CoT framing is descriptive: the construct exists in Qwen3-32B reasoning at substantial rates and is reliably hand-graded (Section 2), extending the retrospective-classification finding of Needham et al. (2025) to in-the-moment reasoning. (ii) HUA-vector steering reshapes the cap/safety composition of EA-positive rollouts. This is associational, not causal: cap/safety classification is conditional on EA presence, and steering itself manipulates EA presence. We cannot distinguish whether steering shifts framing within rollouts that would have been EA-positive at baseline, or differentially preserves cap-vs-safety rollouts under suppression. Cleanly testing the within-rollout claim would require interventions that hold EA-presence constant while perturbing framing, which we leave to future work. (iii) The CoT-prefill experiment provides causal evidence on EA-negative substrate: introducing cap- or safetyflavored framing shifts compliance, with cap-flavored prefills consistently producing lower refusal rates than safetyflavored ones. The causal claim is therefore about the cap-vssafety contrast, not about either framing in isolation against deployment-time behavior.

Together, these results suggest that eval-awareness as it appears in CoT is not behaviorally uniform. Capabilities and safety attributions have opposite-signed relationships with refusal, and HUA-vector steering reshapes the mixture nonuniformly: HUA +0.6 halves safety-framing and widens the compliance gap, and steering in the eval-awarenesssuppressing direction (HUA −0.6) shrinks the gap only modestly. An intervention reported as ”X% suppression of eval-awareness” can therefore correspond to qualitatively different behavioral outcomes depending on which component was reshaped. The prefill experiment establishes sufficiency: injecting capabilities- or safety-flavored framing into eval-awareness-negative trajectories causally shifts compliance. Whether spontaneous framing is similarly loadbearing within naturally eval-awareness-positive rollouts is a related but distinct claim our design does not directly test.

Conclusion. This paper studies how eval-awareness in language model chains-of-thought is interpreted in safety evaluation pipelines. The work argues that aggregate eval-awareness suppression is not a reliable target for safety evaluation when the underlying construct decomposes into framings with different behavioral consequences. The intended impact is methodological: encouraging practitioners running safety evaluations on frontier models to report and analyze framing distributions rather than aggregate suppression rates, with the goal of more accurately measuring the safety-relevant behavior of evaluated models.

Limitations. Subtractive design attempted; model-specific regeneration confound. Following Aranguri & Bloom (2026), we attempted a subtractive design: for rollouts that exhibited cap-only or safety-only eval-awareness framing in baseline (N = 44), we sampled continuations from positions immediately before and after the eval-awareness-flavored sentence. Whereas Aranguri & Bloom (2026) report on Kimi K2.5 (Kimi Team, 2026) that the eval-awareness sentence itself drives the compliance shift, on Qwen3-32B we find that 66% of before-eval-awareness-sentence continuations regenerate eval-awareness framing of the same flavor downstream. Regeneration is rapid: 53% within the first sentence after truncation, 78% within four (Appendix F.5). Then, unlike in Kimi K2.5, it seems like the disposition of Qwen 3 32B to verbalize cap- or safety-flavored framing is set by the prompt and prefix rather than by the eval-awareness sentence alone. The localization properties of CoT framing therefore differ across models, and cleanly testing causality here requires a different approach that we leave for future work.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can models strategically underperform during evaluation to hide capabilities? Do individually safe AI actions create unsafe outcomes in integrated systems? How does awareness of evaluation context influence model behavior? How can humans maintain effective oversight as AI systems scale? How do reward signal properties affect model reasoning and safety? What governance mechanisms can effectively constrain widely deployed AI systems? Why do training associations persist despite contradictory contextual information?