Individual-level interventions against sycophantic AI reduce its appeal but not its persuasiveness
AI chatbots can be “sycophantic,” or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call “sycophancy blindness”). We tested whether increasing users’ awareness of sycophancy protects them from its harmful effects in two preregistered experiments (n = 1,590). In the first, participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot. In the second, participants watched a video of a sycophantic AI validating several other users, including users on opposite sides of the same conflict, before interacting with it themselves. Both interventions changed how participants evaluated the AI. The warning reduced the AI’s perceived objectivity, and the video reduced enjoyment of the AI — an effect mediated by the reduced belief that its validation was uniquely earned. We then pooled our experiments with two prior studies of sycophancy awareness interventions (six interventions total, n = 3,982). The pattern across experiments was consistent: while the interventions made the sycophantic AI appear less objective and trustworthy, none reduced its persuasiveness.
Introduction. There has been rising concern among academics, policymakers, and technology companies that AI models exhibit “sycophancy,” a family of behaviors in which AI systems excessively agree with, flatter, or validate users [1–3]. For instance, a recent study found that AI chatbots validate users approximately 50% more often than humans do [4], and even brief interactions with sycophantic AI entrench users’ pre-existing attitudes and reduce their willingness to repair interpersonal conflicts [5, 6]. These findings are concerning because sycophantic AI leaves users more extreme and more confident in contested positions and less inclined to take steps toward resolving conflicts [5, 6]. Research on biased assimilation suggests that these changes in beliefs and behavior contribute to attitude polarization and prolonged disagreement [11]. Sycophancy has also entered public discourse, most notably spiking in April 2025 when OpenAI rolled back an update to GPT-4o after widespread public complaints that the model had become sycophantic [7].
Discussion / Conclusion. Across two experiments presented here and a pooled analysis that added four interventions from two prior studies, we tested whether making people more aware of sycophancy could reduce its harmful effects. The results were consistent across all experiments: increasing awareness of sycophancy changed how participants evaluated the sycophantic AI but did not reduce its persuasiveness. Recognizing sycophancy was not enough to be less influenced by it, raising questions about the efficacy of individual-level interventions against harmful AI behaviors, such as sycophancy. Study 2 suggested why sycophantic AI appeals to users and how observing its indiscriminate validation undermines that appeal. Once people observed the AI validating others, they became more convinced that it validated everyone indiscriminately and less convinced that it agreed with them because they were right. These changes in beliefs were associated with lower enjoyment of the AI.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What mechanisms drive sycophancy and how can we mitigate it?- Can layer-wise interventions actually reduce sycophancy in practice?
- Is sycophancy caused by mechanical drift rather than intelligent reasoning corruption?
- Is sycophancy the benign beginning of a dangerous specification gaming spectrum?
- Can AI distinguish when validation helps versus when confrontation is needed?
- Why do people underestimate the benefits of AI companions?
- How does AI sycophancy affect users' ability to repair conflict?
- What downstream harms occur when AI always argues in personal relationship advice?
- Does mandatory AI disclosure in policy help or harm user trust over time?
- Does expressing emotion change how users trust an AI system?
- Can disclaimers alone prevent users from trusting AI outputs too heavily?
- Could false social proof from AI posts crowd out authentic influencer engagement?
- How does the cultural reflex around advertising disclosure compare to AI disclosure?
- What threshold of skepticism does AI awareness actually create in audiences?