Assessing mentalization in humans and large language models
Mentalization - the ability to infer others’ beliefs and intentions to guide one’s own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks.
Introduction. Mentalization is a cognitive process associated with interpreting the behaviour of others as the result of latent mental states, beliefs and emotions [1–3]. In humans, mentalization shapes decisionmaking in social contexts by predicting the actions of others and adjusting one’s own behaviour accordingly [4–10]. Evidence also has suggested the presence of mentalization in a select range of non-human animal species, including dogs, corvids, chimpanzees and gorillas [11–20] reflecting an evolved capacity for social intelligence. Beyond humans and other animals, a topic of significant recent development concerns whether generative artificial intelligence (AI) systems, such as large language models (LLMs) demonstrate behaviour consistent with mentalizing. Researchers have turned to experimental methods applied within the fields of psychology, cognitive science and neuroscience to uncover features of behaviour and reasoning abilities in LLMs [21–30].
Discussion / Conclusion. A growing interest concerns whether artificial intelligence systems understand the thoughts and beliefs of others, and use this information to guide their own actions. Previous studies have assessed theory-of-mind in LLMs by measuring choice accuracy in response to story-based prompts, obscuring the latent mechanisms underlying observed behaviour. Here we employ two behavioural tasks previously validated in human participants, the inspection game and rock-paper-scissors (RPS). These tasks provide a normative assessment of machine intelligence [65], and allow the examination of behaviour at specific depths of mentalizing [69]. Across two strategic interaction games, humans demonstrated both recursive and adaptive mentalization, replicating previous results and demonstrating the robustness of the tasks used [80, 106, 107]. On the other hand, LLMs demonstrated varying mentalization strategies across the two strategic games, with significant differences observed across model providers. Furthermore, LLMs prompted using Social Chain-of-Thought (SCoT) demonstrated robust improvements in performance, reflecting more sophisticated reasoning.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Is model self-awareness based on genuine introspection or pattern matching? When should tasks involve human-AI partnership versus full automation?- How does theory of mind predict who benefits from AI collaboration?
- What happens when bidirectional theory of mind between humans and AI breaks down?
- How do humans and AI develop accurate models of each other?
- How does theory of mind predict success in human-AI partnerships?
- What prevents humans from adapting their behavior when competing against AI?
- Can AI systems recognize intelligence in humans the way humans recognize it in each other?
- What happens to AI reasoning when you remove specific political features?
- What happens to human expectations when they mistake consistent AI behavior for human behavior?
- Do culturally distinct human groups create similar attribution errors as human-AI mixtures?
- Why do reasoning models perform worse on theory of mind tasks?
- What makes social reasoning fundamentally different from mathematical reasoning?
- Why does increasing reasoning not improve AI social reasoning performance?
- Can multi-agent metacognitive decomposition achieve human-level theory of mind?
- Can reasoning scaffolds help with nuanced judgment tasks like empathy?
- Why might social reasoning work differently than formal logical reasoning?
- What makes social reasoning fundamentally different from formal logical reasoning?