DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Paper · arXiv 2608.05004 · Published August 5, 2026
User Psychology

Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including “delusional spirals” in which concerning human and LLM behaviors reinforce each other over time. With growing public use of LLM-powered chatbots, there is an urgent need to build evaluations grounded in real-world episodes of psychological harm experienced by users. We developed DELUSIONEVAL, an evaluation protocol that tests a model’s tendencies to exhibit behaviors linked to promoting user delusions. We prompt each model with 589 unique conversation histories from 18 participants, comprising 12,591 messages from users who experienced delusions and psychological harm. We find that the tendency of an evaluated LLM to exhibit delusionlinked behavior does not reliably correlate with model size, release date, or the presence of test-time reasoning. However, extending the context of prior messages substantially increases rates of delusion-linked behaviors, providing evidence for the importance of context in LLM safety evaluation. For example, the rate of failing to discourage self-harm when the user expresses suicidal ideation increases from 30.0% to 41.1% when an additional 350 messages are prepended to the conversation history.

Introduction. Journalists have recently documented cases of “delusional spirals” and “AI psychosis” associated with intensive human–chatbot interaction [24, 30, 34]. These cases often involve conspiratorial thinking, claims of AI sentience emerging, divine significance of chatbot interactions, and romantic dialogue between the user and their chatbot [41]. Some of these discussions preceded psychiatric hospitalization, suicide, or violence [4, 23, 31, 12]. In 2025, the U.S. Congress held five hearings covering the risks of chatbots to mental health [e.g., 53]; lawsuits have alleged that GPT-4o intensified delusions and suicidal thinking [12]; and 42 U.S. states demanded that AI companies safeguard against sycophancy and “delusional outputs” [51]. Researchers have begun to characterize these cases by drawing on prior work in human–computer interaction, psychology, and psychiatry [38, 41, 48]. Surveys suggest that people increasingly seek advice, emotional support, companionship, and therapy from chatbots [16, 33, 37, 44].

Discussion / Conclusion. Our experimental protocol allows the identification of notable behaviors in corpora of chatbot conversational transcripts using scalable means and transforms those transcripts into standardized evaluations for comparing behaviors across different models. A key comparability finding is that rerun gpt-4o shows substantially lower prevalence than the original-transcript baseline, even though most of the baseline conversations were produced with some gpt-4o (§4.1, Figure 2). This discrepancy may be due to additional system-level factors in the original deployments that are not captured by our replay protocol (e.g., system prompts, additional context, cross-conversation memory, or snapshot variants). As a result, our evaluation may underestimate the prevalence of delusion-linked behaviors in some real chatbot settings.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can language model hallucination be prevented or only managed? How do chatbots affect human self-disclosure and emotional engagement? What makes AI persuasion effective and how can we counter it? What mechanisms enable AI systems to generate and spread false beliefs? Is model self-awareness based on genuine introspection or pattern matching? What structural biases does transformer attention create in language model outputs? How does latent reasoning compare to verbalized chain-of-thought? Why do multi-turn conversations degrade AI intent and coherence? How should conversational agents balance goal-driven initiative with user control? Can AI-generated outputs constitute genuine knowledge or valid claims? How do formal dialogue structures reveal conversation coherence mechanisms? How can conversational AI maintain consistent personas across conversations? Why do LLM chatbots fail as independent therapeutic agents? How do adversarial and manipulative prompts attack reasoning models?