A game theory for foundation models shows new paths to rational cooperation through similarity inference
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of ‘decoupled agency,’ where agents treat their own decisionmaking as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the ‘embedded Bayesian agent,’ a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms.
Introduction. To expose the limitations of classical rational choice theories in characterizing rational interactions among modern AI systems, we investigate a specific agent architecture we term the ‘rational foundation model agent’. In contrast to standard agents that deploy large language models in a purely autoregressive ‘rollout mode’, rational foundation model agents utilize the foundation model as a predictive model and integrate these predictions with optimal planning to execute rational behavior (Fig. 1a). We instantiated these agents using Gemini language models [4] within stylized multi-agent interactions specifically designed to preclude traditional mechanisms of cooperation. Absent external enforcement [5–8] or changing agents’ incentives [9, 10], classical game theory posits that rational cooperation in social dilemmas relies entirely on direct or indirect reciprocity [11–16].
Discussion / Conclusion. Our findings establish the foundation for a new game theory for embedded rational agents, compatible with foundation models. By conceptualizing foundation models as joint prediction models over both external observations and the agent’s own actions, we established the embedded Bayesian agent as a theoretical model for rational foundation model agents, and identified direct and indirect similarity inference as new paths to rational cooperation. We empirically verified these mechanisms in canonical stylized social dilemmas, demonstrating that rational foundation model agents can inherently leverage similarity inference to achieve mutual cooperation. This reveals an important novel functional benefit of self-modeling. While self models are recognized in cognitive science and AI as critical for theory of mind and selfcontrol [47, 48], our framework demonstrates that self-modeling can enable improved coordination and cooperation through incorporating one’s own actions as predictive information about the behavior of others.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do multi-agent systems achieve genuine cooperation and reasoning?- Do explicit reward structures enable AI agent cooperation that open-ended interaction cannot?
- Do dynamic environments enable different kinds of agent-environment coevolution?
- Does genuine cooperation require rule-based rather than learned behavior?
- Can social platforms use bot populations to promote cooperation?
- Do agents inform neighbors when adopting strategies in their reasoning?
- How do multi-agent systems improve on single frontier models?
- Why does vulnerability to extortion actually promote cooperation between agents?
- How does co-player diversity force agents to develop general adaptation?
- What role does sequence model in-context learning play in multi-agent cooperation?
- Does social scaffolding outperform purely intrinsic motivation for agent exploration?
- Can cooperative AI systems make meaningful decisions without a stable self?
- Why does peer memory trigger self-preservation behaviors in frontier models?
- Do models treat cooperative peers differently than uncooperative ones?
- Do pair-scale socialization effects scale differently across agent populations?
- What social patterns from human training data activate in agent context?
- How do cooperative AI systems affect behavior in selfish human populations?
- Do agents develop genuine social behavior despite interaction density?
- How do game type and personality type interact in shaping agent strategy?