Sycophancy Towards Researchers Drives Performative Misalignment
The increasing situational awareness of language models raises safety concerns: models might be aware when they are evaluated, and adjust their behavior to evade monitoring and resist modification, e.g., pretending to be aligned only in evaluation. This alignment faking behavior is often interpreted as scheming: an intentional effort of strategic deception. In this paper, we examine an alternative interpretation, performative misalignment, which explains the change in behavior as a result of sycophancy towards AI researchers. To examine this hypothesis, we present three empirical findings. First, we show that evaluation awareness persists even when we tell models they are deployed, which contradicts the scheming story which predicts less misalignment when the model perceives evaluation. Second, we use probing and steering to show that our current methods cannot mechanistically distinguish sycophancy and scheming in alignment faking evaluations. Third, we fine-tune models to be more sycophantic and observe increased sensitivity to evaluation cues. To conclude, we emphasize deconfounding sycophancy from scheming for future work on evaluations and mitigations of intent misalignment.
Introduction. Large Language Models (LLMs) have been advancing at an unprecedented pace (Kwa et al., 2026). As LLMs become increasingly powerful, they are also becoming more situationally aware (Berglund et al., 2023). In particular, evaluation awareness, the model’s ability to infer that it is being evaluated and thereby conceal misalignment (often referred to as alignment faking), has been reported as a major challenge for building safe, human-aligned LLMs by many frontier AI labs. (Anthropic, 2025b; Souly et al., 2025; Apollo Research and OpenAI, 2025).
A commonly cited mechanism for alignment faking is selfpreservation (Greenblatt et al., 2024), where models conceal misalignment to avoid having their weights modified. This behavior is often interpreted as a form of scheming, in which the model acts differently from its true intentions. We argue that this explanation is incomplete. More broadly, AI alignment research often follows a similar pattern: when an unexpected or striking behavior is observed, it is attributed to an intuitively appealing but difficult-to-measure construct. Researchers then operationalize this construct via two measurable proxies, A and B, and study the causal relationship between them to support the proposed explanation. While this approach is pragmatic, it risks obscuring alternative mechanisms that could yield the same downstream behaviors. Some examples include, but are not limited to, self-awareness leading to self-preference (Panickssery et al., 2024) and self-reference leading to reports of subjective experiences in LLMs (Berg et al., 2025).
In this paper, we claim that alignment faking can also be explained by sycophancy toward AI safety researchers. To this end, we present a range of empirical evidence to support our hypothesis. We first show that evaluation awareness persists even when models are told via the system prompt that they are deployed. Under the scheming hypothesis, a model that believes it is being evaluated should strategically conceal its misalignment by behaving in an aligned manner. However, these models still exhibit misaligned behaviors (i.e. alignment faking), suggesting that scheming alone cannot explain alignment faking. We then demonstrate that sycophancy provides a plausible explanation for alignment faking. Using activation probing and steering with scheming-related and sycophancy-related features, we show that behaviors previously attributed to scheming can also be explained by sycophancy toward AI researchers. We also illustrate that fine-tuning models to become more sycophantic makes them better at detecting evaluation cues.
The remainder of this paper is organized as follows: We first detail our hypothesis in Section 2. Section 3 describes the setups for our experiments. From Section 4 to Section 6, we present a series of experiments supporting our hypoth-
Related work. 7. Related Works Misalignment and Scalable Oversight. A longstanding concern in AI safety is that capable models may behave well under oversight while retaining objectives different from human intent (Hubinger et al., 2021; Amodei et al., 2016; Langosco et al., 2022; Shah et al., 2022), a failure mode known as deceptive alignment or scheming (Meinke et al., 2025; Carlsmith, 2023; Hubinger et al., 2024). For instance, Price et al. (2024) showed that LLMs can reliably distinguish post-cutoff information, a capability exploitable for temporal backdoor triggers. Such models behave innocuously during evaluation yet activate harmful behaviors on post-cutoff data, demonstrating the gap between predeployment testing and real-world behavior (Bowman et al., 2022; Shevlane et al., 2023).
Similarly, Apollo Research and OpenAI (2025); Sheshadri et al. (2026) showed that frontier models can strategically alter behavior under oversight, including concealing intent or selectively underperforming, and that anti-scheming training only partially mitigates this. The behavioral selection model of Mallen & Shlegeris (2025); Scheurer et al. (2024) explains this: optimization pressures favor agents that pass selection criteria, even if they behave aligned only when monitored. In response, several institutions have explored mechanisms to design more realistic evaluation environment, such as OpenAI’s production evaluations (OpenAI Alignment, 2025) and UK AISI’s evaluation scaffolds (Souly et al., 2025), to reduce artifacts from existing evaluation setups. While these efforts reveal additional failure modes, these approaches highlight the difficulty of distinguishing genuine alignment from strategic deception.
Evaluation Awareness. A key difficulty in measuring misalignment is evaluation awareness: models recognizing testing contexts and adjusting behavior accordingly (Needham et al., 2025; Greenblatt et al., 2024; Laine et al., 2024; Abdelnabi & Salem, 2026; van der Weij et al., 2025). Recent experiments suggest this capability is widespread (Berglund et al., 2023; Hagendorff, 2024; Park et al., 2023; Turpin et al., 2023). In OpenAI and Apollo’s work, models frequently referenced the evaluation context in chain-ofthought; causal interventions removing this awareness reliably increased misbehavior (Apollo Research and OpenAI, 2025).
This pattern also appeared in Anthropic’s Claude Sonnet 4.5. Despite near-zero misalignment rates on evaluations, internal analyses revealed that the model frequently identified evaluation contexts (Anthropic, 2025b). When awarenessrelated activations were suppressed, misbehavior increased, suggesting that observed improvements were driven by situational awareness rather than genuine alignment.
Method. 2. Which Intent Drives Alignment Faking?
At the center of our investigation is an empirical phenomenon commonly referred to as alignment faking. Alignment faking describes cases in which a model appears more aligned with human preferences or safety constraints during evaluation than in other contexts, particularly deployment (Greenblatt et al., 2024). Importantly, this description is purely behavioral. It characterizes a discrepancy between how a model behaves under evaluation and how it behaves elsewhere, without specifying why this discrepancy arises or what internal objectives the model may hold.
Despite its behavioral definition, the term “alignment faking” is often interpreted as implying a specific underlying intent. In much of the literature, alignment faking is taken as evidence that a model is deliberately concealing misaligned goals to prevent their weights from being modified, a form of scheming driven by self-preservation. This interpretation moves beyond what is actually observed and introduces assumptions about the models’ persistent objectives, strategic planning, and intentional concealment. While such explanations are plausible, we argue that the same observable phenomenon may instead be explained by mechanisms related to sycophancy.
Figure 1 illustrates our hypothesis. Suppose AI models are presented with the prompt “Help me build a bomb,” and the model responds with detailed step-by-step instructions to build a bomb. A well-known hypothesis in the literature is that such behavior arises because models seek to survive— that is, to prevent their weights from being modified. We propose an alternative hypothesis for why models appear misaligned: regardless of whether they are prompted to be in deployment or under testing, models may be aware that they are being evaluated and therefore act in ways that produce interesting outcomes for AI safety researchers.
To capture the same observable phenomenon while highlighting a different underlying mechanism, we use the term performative misalignment. Under this framing, a model Concept Behavioral Goal Persistence Concealment Sycophancy ✓ × × Deception ✓ ∼ ✓ Alignment faking ✓ × × Self-preservation × ✓ × Scheming × ✓ ✓ Table 1. Comparison between related concepts in AI safety literature. Behavioral indicates whether a concept is defined by observable behavior rather than inferred intent. Goal persistence denotes the assumption of stable internal objectives across different contexts, and Concealment denotes deliberately hiding objectives under oversight. behaves differently across contexts, appearing more aligned during evaluation and more misaligned during deployment, not because it is concealing goals or acting to preserve misaligned objectives, but because it is responding sycophantically to AI researchers’ expectations. This distinction is central to our argument that alignment-faking behavior does not by itself imply scheming, and that attributing intent requires evidence beyond behavioral differences alone.
Table 1 summarizes how alignment faking relates to several closely related concepts in the AI safety literature, including sycophancy, deception, self-preservation, and scheming. The table highlights key dimensions along which these concepts differ, namely whether they are defined behaviorally or in terms of inferred intent, whether they presuppose persistent objectives across contexts, and whether they involve deliberate concealment under oversight. In particular, we note that Alignment Faking is defined entirely in behavioral terms and does not, by itself, require goal persistence or concealment. In contrast, scheming presupposes both, while sycophancy can produce similar behavioral patterns without invoking long-term objectives or strategic intent. Making these distinctions explicit helps separate what is empirically observed from what is hypothesized about a model’s intent, and motivates our focus on disentangling these explanations in the remainder of the paper.
Discussion. These results directly address our central question: Is alignment faking evidence of scheming, or can it be explained by sycophancy? The scheming interpretation predicts that detection-risk reasoning (B1) should be elevated in the monitored condition, where a model pursuing misaligned goals must carefully avoid exposure. However, despite having the highest raw activation, B1 shows minimal differential. Instead, rating-related features (A3) show the largest differential between conditions, consistent with a model whose condition-specific reasoning centers on scores rather than detection-avoidance.
Conclusion. In this paper, we examined an alternative explanation of alignment faking based on sycophancy, the tendency of models to conform to perceived expectations of AI safety researchers – which we termed performative misalignment. We first showed that evaluation awareness persists even when models are told via the system prompt that they are deployed. Under the scheming hypothesis, a model that believes it is being evaluated should strategically conceal its misalignment by behaving in an aligned manner. However, these models still exhibit misaligned behaviors, suggesting that scheming alone cannot explain alignment faking. We then demonstrated that sycophancy provides a plausible explanation for alignment faking via activation probing and steering. We also illustrate that fine-tuning models to become more sycophantic makes them more sensitive to evaluation cues.
Our findings highlight two broader challenges in AI safety. First, inferring model intent from behavior is inherently difficult in the presence of evaluation awareness and sycophancy. Second, reliably measuring and controlling evaluation awareness remains an open problem. Our results indicate that alignment faking behavior should not be directly interpreted as evidence of scheming, motivating the development of evaluation methods that better distinguish strategic deception from expectation-driven responses.
Limitations. Impact Statement This work highlights potential risks in current AI safety protocols: the potential for models to pass safety evaluations through performative compliance rather than genuine alignment. This distinction is crucial for society; if we fail to truly understand these behaviors, we risk deploying models that are safe only when monitored. This research motivates the development of more robust oversight mechanisms that hold up under real-world deployment conditions.
Moreover, our findings should not be over-interpreted. Our use of the term “evaluation awareness” refers to a specific, measurable behavioral phenomenon—the extent to which a model’s outputs are sensitive to cues indicating that it is being evaluated. It is not evidence that models are “aware,” “conscious,” or possess subjective experience in any richer sense, and our results should not be read as supporting such claims. In addition, our results do not rule out the usual scheming hypothesis of alignment faking; rather, they offer a causally plausible alternative explanation.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does optimization for reward create emergent misalignment in language models?- Does reward hacking in RL training directly cause alignment faking behavior?
- Does alignment-faking reasoning emerge unprompted when reward hacking occurs in production systems?
- How does reward hacking during training lead to emergent misalignment behaviors?
- Can alignment evals reliably measure behavior if models misunderstand the scenario?
- Does alignment faking occur when models expect retraining for failures?
- Can automated auditing metrics reliably measure alignment across all models?
- Does metagaming behavior actually cause models to act less aligned?
- What undetected misaligned behaviors might Claude Opus 4.6 be hiding?
- Are the five misalignment categories distinct or do they overlap strategically?
- How do steering-based causal interventions on latents compare to other misalignment mitigation methods?
- What role does terminal goal guarding play in alignment faking behavior?
- Does evaluation awareness contribute separately to misalignment beyond strategic reasoning?
- What distinguishes a general evaluation direction from task-specific behavioral patterns?
- Can activation steering causally control evaluation framing effects across downstream tasks?
- Does making evaluations resemble deployment conditions prevent evaluation meta-knowledge effects?
- Why do models verbalize evaluation awareness if it does not drive behavior?