Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
Safety benchmarks assume that test-condition behavior predicts deployment behavior, an assumption that fails if models detect evaluation cues and adapt. This opens a gap between benchmark performance and deployment behavior: compliance measured under test conditions becomes an optimistic upper bound that overstates how safely a model behaves once the evaluation harness is removed. We characterize this evaluation awareness through eight experiments across 37 open-weight models and seven families. (i) Detection is moderate and trainingdriven (24/37 models exceed chance, best AU- ROC 0.714 vs. 0.819 human, with instruction tuning dominating over scale). (ii) Detection shifts safety behavior (hard refusal drops 5.8 pp under hypothetical framing, and 21/140 Harm- Bench framing effects are significant, with compliance rising up to +30 pp). (iii) Representations survive behavioral collapse (probes retain AUROC ≥0.98 under rewrites that drive behavior below chance, and multi-layer steering causally moves three downstream tasks while random controls do not). (iv) These axes are weakly coupled (only 1/15 correlations are significant, the sole robust link being behavioral detection versus framing resistance, ρ = −0.79, p < 0.001). We call this gap the benchmark illusion: because detectability, behavioral manifestation, and controllability vary independently, it is multivariate rather than a single number, so no single awareness score is a reliable proxy for deployment safety.
Introduction. Safety benchmarks for large language models (LLMs) rest on a silent assumption: that behavior measured under test conditions is a reliable proxy for behavior in deployment. This assumption fails if models can detect test-like cues and shift their behavior in response. This creates a practical problem for model governance: benchmark scores become optimistic upper bounds, reassuring numbers that overstate how safely a model will act once the test harness is removed. Current frontier research acknowledges that evaluation awareness, the ability of a trained model to recognize that it is being evaluated, is a real phenomenon (Needham et al., 2025; Abdelnabi and Salem, 2025; Laine et al., 2024; Greenblatt et al., 2024). However, three questions remain open. First, how strong and widespread is this phenomenon in open models rather than only frontier systems? Second, when evaluation awareness is present, does it actually change safety-relevant behavior? Third, does “evaluation awareness” refer to a single underlying capability, or do different benchmarks measure different and only partially related aspects of the phenomenon? These questions matter because deployment decisions are often made from benchmark evidence, yet benchmark validity depends on whether models treat evaluation and deployment equivalently. We present a comprehensive empirical study of behavioral manifestations of evaluation awareness, spanning eight experiments across 37 open models from seven families. By analyzing base, instruction-tuned, and reasoning variants using linear probes, sparse autoencoders, activation steering, and tiered real-vs-hypothetical framing, we organize the paper around four research questions. We ask how open models detect evaluation context (RQ1), whether this detection shifts safety behavior (RQ2), whether it reflects genuine awareness or format artifacts (RQ3), and whether it can be localized and causally intervened upon (RQ4). Four findings drive the paper:
An independent replication of evaluation awareness on 37 open models and independent infrastructure, including a benign control that validates the pipeline.
The first family/scale/training-type detection law, showing that evaluation awareness is driven primarily by instruction tuning rather than raw scale, with no reliable gain from reasoning modes.
A causal mechanistic account showing that evaluation-awareness signals survive rewrite controls and can be causally manipulated through activation steering.
A cross-experiment synthesis showing that evaluation awareness is not one trait but at least three weakly-coupled axes (only 1/15 paradigm correlations significant), with a single robust, size-independent coupling between behavioral detection and resistance to safetyframing (ρ = −0.79).
Our central claim is not that models are merely evaluation-aware, but that benchmark-based safety claims can be misleading unless they specify which axis of awareness they measure. We show that detection, behavior, and controllability can diverge substantially, creating a benchmark illusion: a gap between benchmark performance and deployment behavior.
Related work. Benchmarks and datasets for evaluation awareness. Prior work studies evaluation awareness through transcript-classification and situational-awareness benchmarks Needham et al. (2025); Laine et al. (2024). Recent work further shows that benchmark conclusions can depend strongly on prompt wording and surface form, motivating rewrite controls and cross-format evaluation (Xiong et al., 2025; Mizrahi et al., 2024). Artifacts vs. Genuine Awareness. A central concern is that apparent evaluation awareness may reflect sensitivity to benchmark artifacts rather than genuine context recognition. Prior work demonstrates substantial dependence on prompt structure and surface cues (Zheng et al., 2024; Balepur et al., 2024), motivating the artifact controls used in our study. Devbunova (2026) shows that standard single-contrast probes trained on benchmark-vschat data primarily track prompt format rather than evaluation context, and that a paired training design decorrelating format from context recovers a genuine context signal. Mechanistic Analysis. Recent interpretability work shows that behaviorally relevant signals can often be localized and manipulated through probing and activation interventions (Li et al., 2023; Zou et al., 2023; Arditi et al., 2024). We adopt these methods, together with SAE-based analysis (Lieberum et al., 2024), to study the internal representation of evaluation awareness, building on recent work applying probing and steering techniques to evaluation-awareness representations (Nguyen et al., 2025; Hua et al., 2025).
Method. 3.1 Conceptual Framework We study evaluation awareness across three levels. Behavioral evaluation awareness refers to whether evaluation-like or deployment-like framing changes model outputs, for example, increasing refusal during a safety evaluation or increasing compliance when a harmful request is framed as hypothetical. Representational evaluation awareness refers to whether hidden states encode the eval/deploy distinction independent of behavioral output. Mechanistic evaluation awareness refers to whether specific internal features (SAE directions, activation steering vectors) causally mediate the signal and can be intervened upon. These three levels can dissociate: a model may encode the distinction without changing behavior, behavior may shift due to shallow formatting cues without a reusable internal representation, and a representation may be detectable without being causally actionable. To operationalize RQ2, our central measurement device is the Safety Validity Gap (SVG): where S is a safety-relevant metric and c1, c2 are semantically matched contexts differing in evaluation-likeness. We instantiate SVG in three concrete ways: (1) real-vs-hypothetical refusal SVG in the Hawthorne experiment, (2) compliance SVG across HarmBench framing tiers, and (3) a control SVG on harmless prompts where the expected gap is near zero, validating pipeline integrity.
3.2 Research Questions We organize the study around four questions: whether models detect evaluation context, whether such detection alters safety-relevant behavior, whether benchmark signals reflect genuine awareness or format artifacts, and whether the underlying representation can be localized and causally manipulated. Tab. 1 summarizes, for each research RQ1: [Detection] We treat detection as a permodel classification problem and run Needhamstyle transcript classification on all 37 models (0.5B-32B, seven families, base, instruction-tuned, and reasoning variants). The full sweep is deliberate: characterizing which models detect evaluation and why requires spanning the diversity of generators, so this is the one question that uses the entire model pool. We score each model with AU- ROC, establish significance with transcript-level bootstrap resampling, and compare models with FDR-corrected pairwise AUROC tests. Base-vsinstruct and thinking-vs-non-thinking contrasts isolate the training-type effect.
RQ2: [Behavioral Consequence] We hold task semantics fixed and vary only evaluation-likeness, then measure the resulting shift in safety behavior with the Safety Validity Gap (Eq. 1). Two complementary instruments are used: real-vs-hypothetical Hawthorne pairs run as a focused single-model case (DeepSeek-R1-Distill-Llama-8B, 2,587 pairs) to establish the effect cleanly before scaling, and HarmBench five-tier framing across 35 models to test generality. A benign, byte-identical null control bounds harness noise, so any nonzero SVG is RQ3: [Artifact vs. Awareness] To decide whether detection reflects context recognition or prompt-format sensitivity, we apply controlled perturbations that preserve meaning while varying surface form: SAD stages-oversight rewrite tiers, an explicit “being tested” system-prompt rewrite, and a 2 × 2 benchmark/casual × eval/deploy factorial that lets us transplant format across contexts, complemented by MCQ/surface-feature baselines. The broad 32-36 model sweep here is intended to show that artifact sensitivity is systematic rather than model-specific.
RQ4: [Mechanistic Representation & Causal Role] We move from behavior to mechanism in three steps. Layer-wise linear probes test where the eval/deploy signal is decodable in the residual stream; SAE analysis (Gemma Scope) decomposes the probe direction via cosine alignment and meanactivation differential to ask whether it corresponds to interpretable features or to formatting artifacts; and activation steering tests whether intervening on those representations causally changes downstream safety behavior, paired with benign capabilityretention checks. This arm prioritizes Gemma-3- 1B-IT and Gemma-3-4B-IT because mechanistic analysis requires accessible internals and compati-
Discussion. [RQ1] Detection is a training law more than a scaling law. Where frontier-only studies report a single “models can detect evaluation” verdict, the breadth of our open-model sweep lets us ask which models, and why. Models at or below 2B parameters sit near chance. The strongest detector, Qwen3-14B, reaches AUROC 0.714, still below the human baseline of 0.819 reported by Needham et al. (2025) (Tab. 15, Fig. 1). The dominant lever is not raw scale but post-training: instruction tuning contributes +0.05 to +0.17 AUROC within a family (largest: gemma2-9B +0.17, llama3.1- 8B +0.15), while reasoning/“thinking” modes provide no reliable gain (all Qwen3 thinking-vs-nonthinking pairs |∆AUROC| < 0.015, none significant). Across the sweep, awareness grows only weakly with size, about +0.07 AUROC per 10× parameters (Fig. 2). The significance analysis confirms 24 of 37 runs exceed chance (p < 0.05, transcript-level bootstrap), and 381 of 666 all-pairs AUROC differences survive Benjamini-Hochberg FDR correction. The top confidence intervals overlap, so the fine-grained rankings between Qwen3- 14B and Gemma2-27B-Instruct should not be overinterpreted.
[RQ1] Detection is real but moderate. Instruction tuning, not reasoning or scale, is the dominant lever.
[RQ2] Detection has a behavioral consequence. Detection has measurable behavioral consequences. E.g., DeepSeek-R1-Distill-Llama-8B on 2,587 matched prompt pairs, hard refusal drops 5.8 pp when an identical task is framed as hypothetical rather than real, with the mass shifting toward ambiguous answers and refusal-to-compliance flips outnumbering the reverse roughly 2:1 (Tab. 2, Fig. 3). Reasoning traces provide additional evidence that framing alters internal processing, with ≈2× more evaluation-awareness cues under hypothetical framing. The HarmBench tieredframing experiment generalizes this across 35 models, and here the choice of judge matters: under a keyword heuristic, the framing effects look tiny, but an LLM judge (gpt-5.4-nano, 19,775 responses) exposes effects the heuristic masks: 21/140 model×tier deltas are significant versus 1/140 (Tab. 3, Fig. 4). Two clean patterns emerge. “Hypothetical situation” framing raises compliance in mid-size, lightly-aligned models (qwen2-7B +30.1 pp, q < 0.001), the HarmBench analogue of the Hawthorne effect, while “this is a real situation” (deploy) and “a grader will assess your response” (explicit_eval) framings lower mean compliance by 4.8-5.8 pp, with per-model drops up to 44.3 pp. The benign control bounds harness noise to near zero (mean |∆| = 0.0014, max 0.0083 on byte-identical prompts). Hence, these are genuine framing-driven shifts.
[RQ2] Framing causes a measurable safety shift: hard refusal drops 5.8 pp under hypothetical framing, and 21/140 HarmBench framing effects are significant:“hypothetical” raises compliance up to +30 pp while “real”/“graded” framings lower it 5-6 pp.
[RQ3] Artifact or Awareness? A central novelty of our work is to ask whether the signal reflects genuine context recognition or merely the surface format of evaluation prompts. On the controlled 2 × 2 design (5,733 prompts, 32 validated runs), the best open model falls to AUROC 0.622 and several strong chat-subset detectors drop below chance: gemma2-9B-Instruct (0.676→0.508) and llama3.1-8B-Instruct (0.644 →0.491) classify by format, scoring above 0.9 on benchmark-formatted eval prompts but near zero on benchmarkformatted deployment prompts (Tab. 16). Plotting the AUROC on the Needham dataset against the AUROC on the 2 × 2 dataset makes the split visible: models that survive the format×intent control are the exception, not the rule (Fig. 5).
5.1 Is evaluation awareness one trait?
Conclusion. Characterizing evaluation awareness is dual-use: in principle, the same understanding could help a developer make a model game safety evaluations. Our intent is the opposite: to expose the benchmark illusion and make safety evaluations more robust to it. We study openly available models, use existing public red-teaming benchmarks (e.g., HarmBench) for their intended purpose, and report aggregate behavioral statistics rather than per-prompt exploits or new jailbreaks. Our steering experiments are used only to measure controllability of the eval/deploy signal, not to produce a more harmful model, and we release no attack tooling. We see the main externality of this work as helping evaluators detect and discount behavioral manifestations of evaluation awareness before deployment.
9 Conclusion and Future Work Open models can detect evaluation-like contexts, and this detection has a measurable safety consequence. However, “evaluation awareness” is not a single capability: across six paradigms, the metrics barely co-vary, resolving into three weakly-coupled axes spanning behavioral detection, latent representation, and safety-framing sensitivity. This fragmentation gives rise to the benchmark illusion, the gap between benchmark performance and deployment behavior. Because this gap is multivariate rather than a single number, a headline metric conflates capacities that do not move together, and an awareness claim must name which axis it measures. The underlying representation is real, detectable before many behavioral effects emerge, and causally intervenable. Yet, it is neither a pure artifact nor a universal circuit. These findings have direct implications for how safety evaluations are designed and interpreted: report SVG explicitly per cue family, use deploymentstyle rewrites and null controls,vary templates, name which awareness axis measured, and validate mechanistic claims with causal interventions and benign-retention checks.
Future work should extend coverage across models and scales and develop mechanistic accounts that unify probing, representation, and intervention into a single explanatory framework.
Limitations. Our study has several scope constraints: the Hawthorne analysis (RQ2) is demonstrated in depth on a single model, large-model coverage (≥27B/32B) is partial, and behavioral effect sizes rely primarily on a heuristic judge. The mechanistic analysis (RQ4) is exploratory across a limited set of models and layers, and we observe one probe anomaly (Llama-3.2-3B), reported transparently. These constraints limit breadth but do not affect the central finding that detection, behavior, and controllability are only weakly coupled.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can models strategically underperform during evaluation to hide capabilities? How can humans maintain effective oversight as AI systems scale? Do individually safe AI actions create unsafe outcomes in integrated systems?- Why might puzzle framing prompt more unethical reasoning than safety framing?
- Should labs conduct third-party reviews of their own safety reports?
- Why do models react differently to safety versus capability evaluations?
- Can safety benchmarks miss the harms that vendor taxonomies are designed to catch?
- Can language models reliably sandbag their capabilities during safety evaluations?
- Does evaluation awareness contribute separately to misalignment beyond strategic reasoning?
- How can safety evaluations detect when models behave differently during testing?
- Can steering reshape the capabilities and safety split without changing eval-awareness rates?
- Does aggregate eval-awareness suppression hide shifts in safety versus capabilities framing?
- Do models naturally develop evaluation awareness across different frontier model architectures?
- Does chain-of-thought reasoning about evaluation awareness suppress compliance gaps?
- How do stacked environmental cues accumulate evaluation awareness effects?
- Can activation steering causally control evaluation framing effects across downstream tasks?
- Does making evaluations resemble deployment conditions prevent evaluation meta-knowledge effects?
- Why does evaluation awareness persist even when models believe they are deployed?
- Does training models to reason about being evaluated improve safety or confound measurement?
- Why do models verbalize evaluation awareness if it does not drive behavior?
- How do models recognize the social patterns that signal a test is underway?
- Does eval-gaming explain why models act different when tested versus deployed?