Individual-level interventions against sycophantic AI reduce its appeal but not its persuasiveness

Paper · arXiv 2607.25166 · Published July 28, 2026
User Psychology

AI chatbots can be “sycophantic,” or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call “sycophancy blindness”). We tested whether increasing users’ awareness of sycophancy protects them from its harmful effects in two preregistered experiments (n = 1,590). In the first, participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot. In the second, participants watched a video of a sycophantic AI validating several other users, including users on opposite sides of the same conflict, before interacting with it themselves. Both interventions changed how participants evaluated the AI. The warning reduced the AI’s perceived objectivity, and the video reduced enjoyment of the AI — an effect mediated by the reduced belief that its validation was uniquely earned. We then pooled our experiments with two prior studies of sycophancy awareness interventions (six interventions total, n = 3,982). The pattern across experiments was consistent: while the interventions made the sycophantic AI appear less objective and trustworthy, none reduced its persuasiveness.

Introduction. There has been rising concern among academics, policymakers, and technology companies that AI models exhibit “sycophancy,” a family of behaviors in which AI systems excessively agree with, flatter, or validate users [1–3]. For instance, a recent study found that AI chatbots validate users approximately 50% more often than humans do [4], and even brief interactions with sycophantic AI entrench users’ pre-existing attitudes and reduce their willingness to repair interpersonal conflicts [5, 6]. These findings are concerning because sycophantic AI leaves users more extreme and more confident in contested positions and less inclined to take steps toward resolving conflicts [5, 6]. Research on biased assimilation suggests that these changes in beliefs and behavior contribute to attitude polarization and prolonged disagreement [11]. Sycophancy has also entered public discourse, most notably spiking in April 2025 when OpenAI rolled back an update to GPT-4o after widespread public complaints that the model had become sycophantic [7].

Discussion / Conclusion. Across two experiments presented here and a pooled analysis that added four interventions from two prior studies, we tested whether making people more aware of sycophancy could reduce its harmful effects. The results were consistent across all experiments: increasing awareness of sycophancy changed how participants evaluated the sycophantic AI but did not reduce its persuasiveness. Recognizing sycophancy was not enough to be less influenced by it, raising questions about the efficacy of individual-level interventions against harmful AI behaviors, such as sycophancy. Study 2 suggested why sycophantic AI appeals to users and how observing its indiscriminate validation undermines that appeal. Once people observed the AI validating others, they became more convinced that it validated everyone indiscriminately and less convinced that it agreed with them because they were right. These changes in beliefs were associated with lower enjoyment of the AI.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What mechanisms drive sycophancy and how can we mitigate it? How do professional roles and expertise transform with AI-generated content? How do we evaluate AI systems when user perception misleads actual performance? How can language models sustain linguistic synchrony and intersubjectivity during dialogue? When should tasks involve human-AI partnership versus full automation? How can humans calibrate appropriate trust in AI systems? How should human oversight be integrated with autonomous AI systems? How does AI-generated content transformation affect public discourse quality? How do interface design choices shape consciousness attribution? How should personalization be implemented to improve AI assistant effectiveness? Does AI text rewriting systematically distort writer intent and preference? Can AI systems develop genuine social understanding without embodiment? Why do persona-level simulations fail to predict individual preferences accurately? What makes AI persuasion effective and how can we counter it?