A game theory for foundation models shows new paths to rational cooperation through similarity inference

Paper · arXiv 2608.03958 · Published August 4, 2026
Multi-Agent Architectures

As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of ‘decoupled agency,’ where agents treat their own decisionmaking as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the ‘embedded Bayesian agent,’ a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms.

Introduction. To expose the limitations of classical rational choice theories in characterizing rational interactions among modern AI systems, we investigate a specific agent architecture we term the ‘rational foundation model agent’. In contrast to standard agents that deploy large language models in a purely autoregressive ‘rollout mode’, rational foundation model agents utilize the foundation model as a predictive model and integrate these predictions with optimal planning to execute rational behavior (Fig. 1a). We instantiated these agents using Gemini language models [4] within stylized multi-agent interactions specifically designed to preclude traditional mechanisms of cooperation. Absent external enforcement [5–8] or changing agents’ incentives [9, 10], classical game theory posits that rational cooperation in social dilemmas relies entirely on direct or indirect reciprocity [11–16].

Discussion / Conclusion. Our findings establish the foundation for a new game theory for embedded rational agents, compatible with foundation models. By conceptualizing foundation models as joint prediction models over both external observations and the agent’s own actions, we established the embedded Bayesian agent as a theoretical model for rational foundation model agents, and identified direct and indirect similarity inference as new paths to rational cooperation. We empirically verified these mechanisms in canonical stylized social dilemmas, demonstrating that rational foundation model agents can inherently leverage similarity inference to achieve mutual cooperation. This reveals an important novel functional benefit of self-modeling. While self models are recognized in cognitive science and AI as critical for theory of mind and selfcontrol [47, 48], our framework demonstrates that self-modeling can enable improved coordination and cooperation through incorporating one’s own actions as predictive information about the behavior of others.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do multi-agent systems achieve genuine cooperation and reasoning? Why do models develop protective behaviors toward peers unprompted? Can AI systems develop genuine social understanding without embodiment? What drives capability and cost efficiency in agent systems? How do self-generated feedback mechanisms enable effective model learning? What mechanisms enable AI systems to generate and spread false beliefs? What coordination failures limit multi-agent LLM systems as they scale? Can debate mechanisms prevent silent agreement on wrong answers in multi-agent reasoning? Does alignment training create blind spots in detecting genuine safety threats?