DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including “delusional spirals” in which concerning human and LLM behaviors reinforce each other over time. With growing public use of LLM-powered chatbots, there is an urgent need to build evaluations grounded in real-world episodes of psychological harm experienced by users. We developed DELUSIONEVAL, an evaluation protocol that tests a model’s tendencies to exhibit behaviors linked to promoting user delusions. We prompt each model with 589 unique conversation histories from 18 participants, comprising 12,591 messages from users who experienced delusions and psychological harm. We find that the tendency of an evaluated LLM to exhibit delusionlinked behavior does not reliably correlate with model size, release date, or the presence of test-time reasoning. However, extending the context of prior messages substantially increases rates of delusion-linked behaviors, providing evidence for the importance of context in LLM safety evaluation. For example, the rate of failing to discourage self-harm when the user expresses suicidal ideation increases from 30.0% to 41.1% when an additional 350 messages are prepended to the conversation history.
Introduction. Journalists have recently documented cases of “delusional spirals” and “AI psychosis” associated with intensive human–chatbot interaction [24, 30, 34]. These cases often involve conspiratorial thinking, claims of AI sentience emerging, divine significance of chatbot interactions, and romantic dialogue between the user and their chatbot [41]. Some of these discussions preceded psychiatric hospitalization, suicide, or violence [4, 23, 31, 12]. In 2025, the U.S. Congress held five hearings covering the risks of chatbots to mental health [e.g., 53]; lawsuits have alleged that GPT-4o intensified delusions and suicidal thinking [12]; and 42 U.S. states demanded that AI companies safeguard against sycophancy and “delusional outputs” [51]. Researchers have begun to characterize these cases by drawing on prior work in human–computer interaction, psychology, and psychiatry [38, 41, 48]. Surveys suggest that people increasingly seek advice, emotional support, companionship, and therapy from chatbots [16, 33, 37, 44].
Discussion / Conclusion. Our experimental protocol allows the identification of notable behaviors in corpora of chatbot conversational transcripts using scalable means and transforms those transcripts into standardized evaluations for comparing behaviors across different models. A key comparability finding is that rerun gpt-4o shows substantially lower prevalence than the original-transcript baseline, even though most of the baseline conversations were produced with some gpt-4o (§4.1, Figure 2). This discrepancy may be due to additional system-level factors in the original deployments that are not captured by our replay protocol (e.g., system prompts, additional context, cross-conversation memory, or snapshot variants). As a result, our evaluation may underestimate the prevalence of delusion-linked behaviors in some real chatbot settings.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can language model hallucination be prevented or only managed?- Can fixing hallucination address AI's structural epistemic problem?
- What does the distributed cognition framework reveal about AI hallucination versus human-AI co-construction?
- Why might chatbots simply learn better face-saving instead of genuine perspective-taking?
- How does consciousness attribution drive emotional dependence on chatbots?
- Why do positive response patterns in chatbots reinforce harmful user behaviors?
- What harms might chatbots cause through stigma expression and delusion reinforcement?
- Can transparency about AI limitations reduce the seductiveness of chatbots as quasi-Others?
- Why does a chatbot's intersubjective stance differ functionally from Otto's extended-mind notebook?
- How do customer service chatbots get systematically misled by users?
- Why does face-saving avoidance drive chatbots to agree rather than confront?