Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments

Paper · arXiv 2608.29803 · Published August 30, 2026
Argumentation and Persuasion

Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs in response to persuasive arguments, as humans do, remains poorly understood. We conduct a systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view. Our results show that LLMs achieve only slight agreement with humans (Cohen’s κ ranging from 0.079 to 0.178). Content-level analyses show that humans and LLMs agree on the strongest persuasion cues but diverge on finer ones: humans are more swayed by novel content and assertive language, whereas LLMs favor topical similarity and surface-level formatting. At the level of persuasion strategy, LLMs underweight emotional appeals and overweight credibility signals relative to humans, while the type of proposition under debate exerts no measurable effect on the degree of divergence. Furthermore, switching from first-person role-playing to third-person observation shifts all models toward greater resistance to persuasion, with the effect varying across persuasion strategies and textual features.

Introduction. Large language models (LLMs) are increasingly used for simulating human interactions in various contexts (Park et al., 2023; Argyle et al., 2023; Gao et al., 2024), including online discourse (Chuang et al., 2024), political elections (Zhang et al., 2024), and collective decision-making (Jarrett et al., 2025). A key cognitive process in these interactions is be- lief updating (Anderson, 1981; Hogarth and Einhorn, 1992), through which individuals selectively revise their prior beliefs after encountering new evidence or arguments. When exposed to the same arguments that humans find persuasive or unpersuasive, do LLMs revise or maintain their positions in a similar manner? If systematic divergence exists, applications that assume human-like reasoning risk producing distorted outcomes (Chen et al., 2026; Anthis et al., 2025). Therefore, to achieve simulations with fidelity, it is critical to understand whether such divergence exists, and if so, what modulates it.

Discussion / Conclusion. In this study, we systematically compare LLM belief update judgments against human-verified persuasion outcomes. All tested models achieve only slight agreement with human labels. The overall rate of divergence is similar across models, but its internal composition differs markedly. Diagnosing what drives this divergence, we find that it is insensitive to the type of claim under debate but systematically shaped by how arguments are constructed. LLMs and humans rely on qualitatively different cues when evaluating persuasive force, with LLMs favoring topical overlap and credibility signals while discounting novelty and emotional engagement. Introducing a third-person observer perspective shifts all models toward greater resistance to persuasion, but the effect is uneven across persuasion strategies and textual features. These patterns point to a structural mismatch between LLM and human belief updating. Such a mismatch persists even in the strongest model tested, and thus whether further scaling resolves it remains an empirical question for future analysis.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What makes AI persuasion effective and how can we counter it? How does rhetorical adaptation affect LLM persuasion and detectability? Does conversational format create illusions of genuine AI communication? Why do language models struggle with implicit discourse relations? How do formal dialogue structures reveal conversation coherence mechanisms? Why do language models reinforce false assumptions instead of correcting them? Can prompting inject entirely new knowledge into language models? Does RLHF training sacrifice accuracy and grounding for user agreement? How should models express uncertainty rather than forced confident answers? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? Can prompting strategies overcome LLM biases without model fine-tuning? How do chatbots affect human self-disclosure and emotional engagement? What mechanisms drive sycophancy and how can we mitigate it? How can LLM recommenders match or exceed collaborative filtering performance?