Does reasoning training improve theory of mind or just stability?
Reasoning models score better on theory of mind tests, but is this a new ability or just more reliable access to existing skills? Understanding what drives the gains matters for knowing whether models actually understand minds better.
The paper's central claim is a reinterpretation of why reasoning models score better on Theory of Mind tests. Using "novel adaptations of machine psychological experiments together with results from established benchmarks," the authors report that reasoning models "consistently exhibit increased robustness to prompt variations and task perturbations." Their analysis suggests the gains come "at least partly" from models being "more robust at reaching the correct answer under prompt and task variation," and they read this "as evidence for a robustness-based account rather than for a new ToM-specific ability." The discussion also says models "have improved substantially in ToM tasks since 2023," and that part of this is "likely just that models have become more capable overall." Robustness is offered as a second factor, not the only one.
The mechanism is stated as a hypothesis rather than a demonstrated result. The authors borrow a finding from Yue et al., that the reasoning paths of reasoning models stay bounded by their base models and that reasoning training "does not appear to add fundamentally new capabilities." Read alongside their own results, this yields the claim that "the main effect of this kind of reasoning training is improved stability in reaching a solution the model could already, in principle, reach, rather than an expansion of representational capacity." The clearest support in the excerpt is the thinking-on version of Claude, which was more robust to prompt and task variation than its thinking-off counterpart. On this view, thinking makes a latent ToM skill more reliable, not larger.
This sits in an unusual place among the neighbors. Two existing notes report that reasoning training does not help social cognition: Why do reasoning models fail at theory of mind tasks? finds regression on Decrypto, and Why do reasoning models struggle with theory of mind tasks? finds no link between reasoning effort and accuracy. This paper agrees that reasoning training adds no ToM-specific mechanism, but it does not agree that nothing improves: it reports better performance since 2023 and a real robustness effect. One way to hold both is that reasoning may steady an existing ability without extending it, so a gain shows up under perturbation more clearly than on a fixed benchmark. That reconciliation is my reading, and the excerpt does not make it. The robustness account also bears on Can language models solve ToM benchmarks without real reasoning? and on Do large language models genuinely simulate mental states?. All three ask what a ToM score measures, and this paper's answer is that it can partly measure how stable a model is under rewording.
The excerpt leaves a lot unestablished. It names no benchmarks, model list, sample sizes or effect sizes beyond the Claude thinking-on versus thinking-off contrast, and the "at least partly" wording leaves open how much of the gain robustness explains. The authors state that the setup "compares thinking and non-thinking models rather than RLVR-trained and non-RLVR-trained ones," so the effect cannot be attributed to RLVR specifically. The bounded-by-base-model premise comes from another paper and is used here as supporting context, not as something this study tests. What follows at this strength is a caution about how to describe ToM results: a reasoning model passing a perturbed ToM task is evidence of reliability, and the excerpt gives no ground for calling it a new capacity to model minds.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why doesn't reasoning volume improve theory of mind performance?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do reasoning models fail at theory of mind tasks?
Recent LLMs optimized for formal reasoning dramatically underperform at social reasoning tasks like false belief and recursive belief modeling. This explores whether reasoning optimization actively degrades the ability to track other agents' mental states.
contrast: Decrypto shows regression for reasoning models, while this paper reports improvement since 2023 and robustness gains
-
Why do reasoning models struggle with theory of mind tasks?
Extended reasoning training helps with math and coding but not social cognition. We explore whether reasoning models can track mental states the way they solve formal problems, and what that reveals about the structure of social reasoning.
agrees that reasoning training adds no ToM-specific mechanism, though this paper finds a robustness effect rather than none
-
Can language models solve ToM benchmarks without real reasoning?
Do current theory-of-mind benchmarks actually measure mental state reasoning, or can models exploit surface patterns and distribution biases to achieve high scores? This matters because it determines whether benchmark performance indicates genuine understanding.
both question what ToM scores measure; this paper points to stability under perturbation as part of the answer
-
Do large language models genuinely simulate mental states?
This explores whether LLMs perform authentic theory of mind reasoning or rely on surface-level pattern matching. The distinction matters because evaluation format—multiple-choice versus open-ended—reveals very different capability levels.
related concern about validity of ToM evidence; this paper adds perturbation robustness as a probe
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
- A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks
- Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models
- Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
- PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
Original note title
the theory of mind gains of reasoning models read as robustness to prompt and task variation rather than a new ability specific to theory of mind