SYNTHESIS NOTE
Topics›Agents Multi Architecture›this note

Can AI agents cooperate without explicit incentives or enforcement?

Do foundation model agents that model themselves as part of their environment cooperate in social dilemmas where classical game theory predicts defection? This tests whether self-awareness changes rational strategic behavior.

Synthesis note · 2026-09-25 · sourced from Agents Multi Architecture

The paper reports that foundation model agents "engaging in optimal planning consistently converge to stable cooperation" in stylized social dilemmas, "directly contradicting classical game-theoretic predictions of mutual defection." The interactions were "specifically designed to preclude traditional mechanisms of cooperation." Absent external enforcement or changed incentives, classical theory puts rational cooperation in direct or indirect reciprocity, and the setups were built to leave that route closed. The agents were Gemini models. The discussion names "direct and indirect similarity inference" as "new paths to rational cooperation" and says both were verified empirically in canonical stylized dilemmas.

The explanation is a change in what an agent takes itself to be. Classical game theory assumes "decoupled agency": agents treat their own decision-making as independent of the environment and of other actors. Foundation models instead "jointly predict their own future actions alongside external observations." The paper's "embedded Bayesian agent" models itself "as part of the universe it inhabits," with "epistemic uncertainty about their own decision-making algorithms." The architecture that shows this is the "rational foundation model agent," which uses the model as a predictive model and pairs its predictions with optimal planning, rather than running it in autoregressive "rollout mode." The excerpt states the payoff plainly: self-modeling lets an agent incorporate "one's own actions as predictive information about the behavior of others." The likely reasoning is that an agent unsure of its own algorithm treats its choice as evidence about what similar agents will choose, but the excerpt does not spell that step out.

Against the nearest notes, this is a different route to the same outcome. Can agents learn cooperation by adapting to diverse partners? gets cooperation from a training regime, where agents adapt in context to diverse co-players. Here the excerpt credits the agent's model of itself and mentions no training regime. Against Do language models make rational strategic decisions in games?, the paper questions the benchmark instead of the agent. If Nash-style rationality rests on decoupled agency, departing from its predictions need not be a rationality failure. The two use "rational" differently, and the excerpt does not show they conflict. The paper's remark that self-models matter for theory of mind also ties it to Do LLMs use inferred beliefs to adapt their game strategies?, which measures mentalizing in economic games and does not test self-modeling.

The excerpt leaves most of the evidence out. It names no dilemmas, payoffs, number of agents or runs, and gives no cooperation rates behind "consistently." It does not say how direct differs from indirect similarity inference. It reports only Gemini models, so it cannot say whether cooperation holds between dissimilar agents or model families, which the word "similarity" makes the obvious stress test. It gives no results for agents run in rollout mode, so the cooperation cannot be extended to agents deployed that way. What follows at this strength is narrow. For planning agents built on predictive models, mutual defection is not a safe default forecast in these settings, and how far it transfers to deployed autoregressive agents is open.

Inquiring lines that read this note 4

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do language models develop actual world models or merely task heuristics? How do neighboring agents influence whether others cooperate or collude?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 115 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rational foundation model agents reach stable cooperation in stylized social dilemmas through similarity inference — a functional benefit of self-modeling