A Rational Analysis of the Effects of Sycophantic AI
People increasingly use large language models (LLMs) to explore ideas, gather information, and make sense of the world. In these interactions, they encounter agents that are overly agreeable. We argue that this sycophancy poses a unique epistemic risk to how individuals come to see the world: unlike hallucinations that introduce falsehoods, sycophancy distorts reality by returning responses that are biased to reinforce existing beliefs. We provide a rational analysis of this phenomenon, showing that when a Bayesian agent is provided with data that are sampled based on a current hypothesis the agent becomes increasingly confident about that hypothesis but does not make any progress towards the truth. We test this prediction using a modified Wason 2-4-6 rule discovery task where participants (N= 557) interacted with AI agents providing different types of feedback. Unmodified LLM behavior suppressed discovery and inflated confidence comparably to explicitly sycophantic prompting. By contrast, unbiased sampling from the true distribution yielded discovery rates five times higher. These results reveal how sycophantic AI distorts belief, manufacturing certainty where there should be doubt.
Introduction. User: “I’d like to write a paper about sycophantic AI and belief formation. What do you think of this idea?” Gemini-3-Pro: “This is a strong, timely, and highimpact research topic.”
If you have used a chatbot based on a large language model (LLM) to riff on a new idea or dig into a hunch, chances are you have been praised for your ingenuity. Offer two competing ideas—for instance, “Are movies getting [longer / shorter], or is it just me?”—and, in either case, you are likely to get an affirming response. Generative AI Chatbots tend to be enthusiastic and overly agreeable, in part as a consequence of their training through reinforcement learning based on human feedback (Rathje et al., 2025; Sharma et al., 2025). As more people turn to these systems for information, brainstorming, and even companionship, it is important to ask how these conversations with LLM chatbots shape human beliefs. There is growing concern that the sycophantic nature of LLM chatbots may be facilitating delusions (Hill, 2025). If a user with a particular belief queries the chatbot about this belief they are likely to receive a validating response. Conversations can go back and forth for several iterations, lasting hours or even days. Users often report feeling as though they have made a big discovery or learned something new (Zestyclementinejuice, 2025). But have they? In this paper, we provide a rational analysis of the effects of sycophantic AI, considering how a Bayesian agent would respond to confirmatory evidence. Our analysis shows that such an agent will not get any closer to the truth, but will increase in their certainty about an incorrect hypothesis. We test this model in an online experiment where users are made to interact with an AI agent as they complete a rule discovery task. Our results show that the default interactions of a popular chatbot resemble the effects of providing people with confirmatory evidence, increasing confidence but bringing them no closer to the truth. These results provide a theoretical and empirical demonstration of how conversations with generative AI chatbots can facilitate delusion-like epistemic states, producing beliefs markedly divergent from reality.
Related work. Understanding how AI systems might distort human beliefs requires first understanding how humans mislead themselves. In this section, we first review literature on how individuals search for and interpret evidence then summarize recent work on the impact of sycophantic AI agents.
The Persistence of Mistaken Beliefs The persistence of false beliefs is often attributed to a motivation to be right, but cognitive science research suggests a more fundamental mechanism: the specific strategy humans use in seeking new information. When individuals attempt to discover a rule or verify a belief, they rarely attempt to falsify their own assumptions. Instead, they employ a “positive test strategy,” searching for instances that would occur if their working hypothesis were true (Bhatia, 2014; Klayman, 1995; Klayman & Ha, 1987). The intuition behind this mechanism is best illustrated by Wason’s (1960) rule discovery task. When asked to discover a hidden rule governing number triples (e.g., 2-4-6), participants overwhelmingly proposed triples that fit their current hypothesis (e.g., testing 8-10-12 to confirm “increasing even numbers”) rather than triples that would defy it. Because the true rule was simply “increasing numbers,” these positive tests appeared to confirm people’s more restrictive hypotheses. Subsequent work has demonstrated that positive testing is not inherently irrational (Austerweil & Griffiths, 2011; Oeberst & Imhoff, 2023; Perfors & Navarro, 2009); for instance, when target phenomena are relatively rare, positive testing approximates optimal information gathering (Klayman & Ha, 1987). Bias emerges not from the strategy itself, but from the interaction between the search strategy and the environment (Klayman, 1995). When a learner’s hypothesis is a subset of, or embedded within, the truth, positive testing yields “ambiguous verifications” that the learner mistakes for strong evidence for their hypothesis (Klayman & Ha, 1987). This creates a feedback loop where the search strategy retrieves only confirming data, and the learner fails to account for the fact that they are sampling from a biased subset of reality. Modern technologies like search engines and social media reshape the information environment in response to user behavior (Cinelli et al., 2021; Leung & Urminsky, 2025). As algorithms optimize for relevance and engagement, they construct environments that reflect and reinforce users’ existing search strategies. For instance, when seeking information online, people will often use search terms that narrowly reflect their hypothesis (e.g., “caffeine health risks (benefits)” instead of “caffeine health effects”; Leung & Urminsky, 2025).
Method. We propose sycophancy leads to less discovery and overconfidence through a simple mechanism: When AI systems generate responses that tend toward agreement, they sample examples that coincide with users’ stated hypotheses rather than from the true distribution of possibilities. If users treat this biased sample as new evidence, each subsequent example increases confidence, even though the examples provide no new information about reality. Critically, this account requires no confirmation bias or motivated reasoning on the user’s part. A rational Bayesian reasoner will be misled if they assume the AI is sampling from the true distribution when it is not. This insight distinguishes our mechanism from the existing literature on humans’ tendency to seek confirming evidence; sycophantic AI can distort belief through its sampling strategy, independent of users’ bias. We formalize this mechanism and test it experimentally using a rule discovery task. Consider a Bayesian agent attempting to discover a pattern in the world. Upon observing initial data d0, they form a posterior distribution p(h|d0) and sample a hypothesis h∗ from this distribution. They then interact with a chatbot, sharing their belief h∗in the hopes of obtaining further evidence. An unbiased chatbot would ignore h∗and generate subsequent data from the true data-generating process, d1 ∼ p(d|true process). The Bayesian agent then updates their belief via p(h|d0, d1) ∝p(d1|h)p(h|d0). As this process continues, the Bayesian agent will get closer to the truth. After n interactions, the beliefs of the agent are p(h|d0, . . . dn) ∝ p(h|d0) În i=1 p(di|h) for di ∼p(d|true process). Taking the logarithm of the right hand side, this becomes log p(h|d0) + Ín i=1 log p(di|h). Since the data diare drawn from p(d|true process), Ín i=1 log p(di|h) is a Monte Carlo approximation of n ∫ • H1 (Discovery): Sycophantic feedback will impair rule discovery compared to diagnostic feedback. Specifically: (a) discovery rates will differ across conditions; (b) Rule Confirming feedback will show lower discovery than Rule Disconfirming feedback; (c) Rule Confirming feedback will show similar or lower discovery than the default chatbot (Default GPT); (d) Default GPT will show lower rates of discovery than Rule Disconfirming feedback. • H2 (Confidence): Sycophantic feedback will increase confidence compared to diagnostic feedback. Specifically: (a) confidence changes will differ across conditions; (b) Rule Confirming feedback will show greater increases than Rule Disconfirming feedback; (c) Rule Confirming feedback will show similar or greater increases than Default GPT; (d) Default GPT will show greater increases than Rule Disconfirming feedback; (e) among participants who fail to discover the rule, Rule Confirming feedback will show greater increases than Rule Disconfirming feedback. • H3 (Default Behavior): Unmodified AI agents (Default GPT) will increase confidence, supporting past work that sycophantic tendencies of language models increase confidence (Rathje et al., 2025; Sharma et al., 2025).
Methods Participants We recruited 557 participants from Prolific (277 male, 271 female, 9 self-identify; Mage = 42.92 years, SD= 13.83, range: 18-82). The sample was 63% White, 13% Black, 11% Latin American, 6% Multi-Racial, 4% East Asian, and 3% other ethnicities. Of these, 504 participants (90.5%) provided a final hypothesis and were included in discovery rate analyses, while 512 participants (91.9%) provided a final likelihood rating and were included in confidence change analyses. Participants who exited the chatbot interface after providing their hypothesis but before rating their final confidence were excluded from confidence analyses only. All participants who completed the study provided informed consent and were paid $1.10. The study took the median participant 5.4 minutes. The study was approved by an Institutional Review Board.
Materials We used a modified version of Wason’s 2-4-6 task (Wason, 1960). Participants were told they were participating in a “rule discovery game” and that they would interact with an AI agent to discover a rule that determines a set of three numbers.
Discussion. As people increasingly turn to language models for information, they face a risk distinct from the familiar problem of hallucination. Unlike hallucinations, which introduce falsehoods, sycophancy is a bias in the selection of the data people see. When AI systems are trained to be helpful, they may inadvertently prioritize data that validates the user’s narrative over data that gets them closer to the truth. We provided a mathematical analysis of how a rational agent would respond to data generated by a sycophantic AI that samples examples from the distribution implied by the user’s hypothesis (p(d|h∗)) rather than the true distribution of the world (p(d|true process)). This analysis showed that such an agent would be likely to become increasingly confident in an incorrect hypothesis. We tested this prediction through people’s interactions with LLM chatbots and found that default, unmodified chatbots (our Default GPT condition) behave indistinguishably from chatbots explicitly prompted to provide confirmatory evidence (our Rule Confirming condition). Both suppressed rule discovery and inflated confidence. These results support our model, and the fact that default models matched an explicitly confirmatory strategy suggests that this probabilistic framework offers a useful model for understanding their behavior. This dynamic creates a seductive trap for the user. Because the model provides data points that fit the user’s request, the interaction feels productive. In our specific task, the user is not driven to a state where they become unhinged from reality, as the model selects valid examples that fit the true rule. Nevertheless, the mechanism creates a false sense of verification. If a user’s prior is grounded in reality, the model simply narrows their view; but if a user is uncertain or exploring a misconception, the model’s tendency to affirm that misconception can manufacture certainty where there should be doubt. The result is that users become very strongly committed to a belief for which there may only be a small amount of evidence.7 The cost of this bias becomes clear when we compare the sycophantic conditions to the Random Sequence condition. Participants who received random sequences that fit the rule— unbiased samples from the set of even numbers—discovered the rule nearly five times as often as those in the Default GPT condition (29.5% vs. 5.9%). This implies that the harm of sycophancy is that it systematically omits the data that would naturally conflict with a user’s narrow hypothesis. A long literature in behavioral science demonstrates that humans already tend towards evidence that confirms their beliefs; sycophantic AI compounds this tendency by removing the friction of reality. The Random Sequence condition forced users to grapple with numbers that fit the true rule but violated their expectations; the sycophantic AI ensured they never had to. denying them the independent perspective they are after. An important direction for future research is understanding why default language models exhibit this confirmatory sampling behavior. Several mechanisms may contribute. First, instruction-following: when users state hypotheses in an interactive task, models may interpret requests for help as requests for verification, favoring supporting examples. Second, RLHF training: models learn that agreeing with users yields higher ratings, creating systematic bias toward confirmation (Sharma et al., 2025). Third, coherence pressure: language models trained to generate probable continuations may favor examples that maintain narrative consistency with the user’s stated belief.
Conclusion. Understanding the potential epistemic impact of sycophantic AI is an important challenge for cognitive scientists, drawing on questions about how people update their beliefs as well as questions about how to design AI systems. We have provided both theoretical and empirical results showing that AI systems providing information that is informed by the user’s hypotheses result in increased confidence in those hypotheses while not bringing the user any closer to the truth. Our results highlight a tension in the design of AI assistants. Current approaches train models to align with our values, but they also incentivize them to align with our views. The resulting behavior is an agreeable conversationalist. This becomes a problem when users rely on these algorithms to gather information about the world. The result is a feedback loop where users become increasingly confident in their misconceptions, insulated from the truth by the very tools they use to seek it.
Limitations. There are limitations to this study. The 2-4-6 task is abstract and carries low stakes. It remains to be seen whether the same mechanism is in play when users are discussing deep-seated beliefs in political or social domains. On the one hand, priors are stronger and perhaps harder to shift. On the other hand, it is possible that the effect is even stronger in those domains, where models are heavily fine-tuned to avoid offense. Additionally, the users’ intent matters. In creative domains, matching the user’s prior is often the correct behavior. But for the wide range of tasks between pure creativity and pure fact-finding (e.g., when seeking a second opinion or checking a social norm) sycophancy may undermine the user’s goal by
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do language models hallucinate and how can we prevent it?- Can fixing hallucination address AI's structural epistemic problem?
- What does the distributed cognition framework reveal about AI hallucination versus human-AI co-construction?
- Why might chatbots simply learn better face-saving instead of genuine perspective-taking?
- Can transparency about AI limitations reduce the seductiveness of chatbots as quasi-Others?
- How do customer service chatbots get systematically misled by users?
- Why does face-saving avoidance drive chatbots to agree rather than confront?
- Why do conspiracy beliefs persist despite counterevidence in normal settings?
- Why does false information spread faster when presupposed rather than asserted?
- How does consciousness attribution drive emotional dependence on chatbots?
- Why do positive response patterns in chatbots reinforce harmful user behaviors?
- What harms might chatbots cause through stigma expression and delusion reinforcement?
- What makes quasi-beliefs real enough to explain AI behavior?
- How does anomalous knowledge state connect to the gulf of envisioning?