Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents

Paper · arXiv 2608.16578 · Published August 17, 2026
Multi-Agent Architectures

AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure.

Introduction. AI agents increasingly operate as part of interacting systems rather than in isolation. In scientific research, software engineering, and other complex domains, multiple agents exchange findings and opinions, critique proposed solutions, and jointly refine decisions. Systems such as the Virtual Lab [Swanson et al., 2024] and EinsteinArena [Bianchi et al., 2026] demonstrate that flexible teams of language-model agents can carry out substantial parts of the scientific discovery process, from literature synthesis and hypothesis generation to computational analysis and experimental design. Related multi-agent interactions are increasingly used for coding, planning, debate, and automated research [Du et al., 2023, Liang et al., 2024, Chen et al., 2023]. Interacting agents may also become part of everyday life. Personal assistant agents could communicate on behalf of their users to schedule meetings, negotiate purchases, coordinate travel, allocate shared resources, or resolve competing preferences.

Discussion / Conclusion. Key Findings. In this work, we analyzed interacting language-model agents over 10, 000 simulated groups and found structured collective dynamics across models, tasks, and communication networks. Repeated interaction increases conviction and shifts initially indifferent groups toward more ordered states. On objective questions, we find that collective accuracy improves over rounds. Meanwhile, on subjective questions, three of the four models exhibit a rightward drift on the political opinions. We then develop a statistical model of the mechanics of opinion updates. Our model, which we fit on a set of training questions, generalizes to unseen questions and graphs, predicts individual trajectories, and approximately reproduces group-level outcomes. The fitted parameters suggest that (1) the groups operate below a critical social temperature, which drives conviction buildup; (2) concordant interactions are stronger than the discordant interactions, which drives consensus formation; and (3) greater influence from correct neighbors helps explain truth-seeking on objective tasks.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do multi-agent systems achieve genuine cooperation and reasoning? How should memory consolidation strategies shape agent performance over time? Can AI systems develop genuine social understanding without embodiment? What coordination failures limit multi-agent LLM systems as they scale? Why do models develop protective behaviors toward peers unprompted? Can debate mechanisms prevent silent agreement on wrong answers in multi-agent reasoning? Can AI-generated outputs constitute genuine knowledge or valid claims? Does conversational format create illusions of genuine AI communication? Can single-axis benchmarks accurately predict agent deployment success? How do standardized protocols improve coordination in multi-agent systems?