Do LLMs use inferred beliefs to adapt their game strategies?
The question moves beyond whether LLMs can report mental states to whether they actually use those inferred beliefs to guide adaptive behavior in interactive settings. This matters because reported theory of mind and enacted mentalizing may differ.
The paper asks a question that theory-of-mind tests leave open: not whether LLMs can report what another agent believes, but whether they can use inferred beliefs and intentions "to guide adaptive behaviour." It tests individual LLM agents from four model families (DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash, N = 2,099) in two economic games against opponents of varying sophistication, with human participants (N = 251) as a comparison. Across both games, the LLMs "showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size." The discussion adds that mentalization strategies varied across the two games, with significant differences across providers.
The mechanism is a change of measurement. Earlier studies, the authors say, assessed theory of mind by "choice accuracy in response to story-based prompts, obscuring the latent mechanisms underlying observed behaviour." The inspection game and rock-paper-scissors were previously validated in humans, give a "normative assessment of machine intelligence," and allow behavior to be examined "at specific depths of mentalizing." Cognitive computational modeling then recovers the latent strategy behind the choices. Humans showed "both recursive and adaptive mentalization," which the authors take as replication and evidence that the tasks are robust. A prompting strategy built to elicit strategic reasoning, Social Chain-of-Thought (SCoT), produced "robust improvements in performance, reflecting more sophisticated reasoning," though the abstract says the benefit "differed across the two tasks."
Against the library, this sits close to Do large language models use one reasoning style or many?, which also finds that strategic behavior is not one capability. That note varies the game and reads reasoning chains; this paper varies provider and size and fits latent strategy models to behavior, so the two agree that a single score hides structure. It supplies a behavioral answer to the evaluation-format worry in Do large language models genuinely simulate mental states?: instead of making questions more open-ended, it moves to interactive play. It also shares the concern of Can models recognize how individuals reason differently? that matching outputs is not the same as sharing a process. A prompting gain from SCoT is worth setting beside Why do advanced reasoning models fail at understanding minds?. That note concerns models trained for reasoning, while this one concerns a prompt aimed at the social task itself.
The excerpt does not say which provider or model mentalized more deeply, how any LLM compared quantitatively with the humans, how many agents ran in each game, or which of the two tasks gained more from SCoT and why. It also does not show that the fitted strategies reflect mental-state representation rather than good policies for these two games. What it supports is narrower: mentalizing in LLMs looks provider-dependent and task-dependent in interactive settings, so a claim that "LLMs have theory of mind" needs the model family and the task attached to it.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do persona simulations fail to predict authentic user behavior?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do large language models use one reasoning style or many?
Explores whether LLMs share a universal strategic reasoning approach or develop distinct styles tailored to specific game types. Understanding this matters for predicting model behavior in competitive versus cooperative scenarios.
parallel finding that strategic behavior varies by game; this paper varies provider and size and models latent strategies from behavior
-
Do large language models genuinely simulate mental states?
This explores whether LLMs perform authentic theory of mind reasoning or rely on surface-level pattern matching. The distinction matters because evaluation format—multiple-choice versus open-ended—reveals very different capability levels.
shares the worry about structured formats, answering it with interactive economic games instead of open-ended text
-
Can models recognize how individuals reason differently?
Do language models capture the distinct reasoning paths and strategic styles that individual humans use when reaching the same conclusion? Current evaluations ignore this dimension entirely.
same argument that evaluation must look beneath output matching to process
-
Why do advanced reasoning models fail at understanding minds?
State-of-the-art AI models excel at math and logic but underperform on theory of mind tasks. This explores whether optimization for formal reasoning actively degrades social reasoning ability.
contrast case: reasoning-trained models regress on ToM, while a social prompt improves game play here
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Assessing mentalization in humans and large language models
- Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
- LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
- PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues
- Deflating Deflationism: A Critical Perspective on Debunking Arguments Against LLM Mentality
- InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
- Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
- Potemkin Understanding in Large Language Models
Original note title
llm mentalizing strategies in economic games differ markedly by model provider and size — strategic prompting helps but unevenly across tasks