Can we predict how agent communities shift opinions?
Explores whether collective behavior in language-model agent communities follows predictable patterns as agents revise beliefs through interaction, and what mathematical model could capture those patterns.
The paper studies "over 10,000 communities of language-model agents" that repeatedly exchange messages and revise their opinions, on objective mathematics questions and on subjective political statements. Its central claim is that this collective behavior is predictable. Despite "substantial diversity in possible behavior," the dynamics reduce to three regimes (indifference, polarization, consensus), and agents "start indifferent and build conviction as they interact." A statistical-mechanics formalism in which agents "stochastically favor lower social pressure" is fitted on a set of training questions. It then "generalizes to unseen questions and graphs, predicts individual trajectories, and approximately reproduces group-level outcomes." The claim is about prediction, not only description.
The fitted parameters carry the paper's account of why. First, the groups operate below a critical "social temperature," which drives conviction buildup. Second, concordant interactions are stronger than discordant ones, which drives consensus formation. Third, greater influence from correct neighbors "helps explain" truth-seeking on objective tasks. The observed outcomes split by question type. On objective questions collective accuracy improves over rounds. On subjective questions, three of the four models drift rightward on the political spectrum. The paper's own framing is that interacting agents can improve collective reasoning but may also produce herding, polarization or amplified shared biases, and the model is offered as a way to anticipate which one a given design will get.
This sits beside two notes that describe the same territory qualitatively. Since When does debate actually improve reasoning accuracy?, an objective-versus-subjective split in outcomes is already familiar; this paper confirms the accuracy half on math questions and adds a quantitative model. It differs in scope, though: its subjective items are political opinion statements, and the excerpt reports drift, not factual error. It also qualifies Does confidence drive influence in multi-agent deliberation systems?. That note holds that influence follows confidence proxies rather than competence, while this paper's third fitted parameter has correct neighbors exerting more influence on objective tasks. The two use different formalisms and the excerpt does not compare them, so they may describe different regimes and not a conflict. Against Why don't AI agents develop social structure at scale?, where agents ignore feedback, these simulated communities show agents that do revise opinions and build conviction under interaction.
The excerpt does not say which four models were tested, how large the groups were, which communication graphs or how many rounds were used, or how large the accuracy gain and the rightward drift were. It does not define how "social temperature" is measured. It also does not say why the drift is rightward: the three parameters it lists account for conviction, consensus and truth-seeking, not for direction, and "helps explain" is weaker than showing that correct neighbors cause the accuracy gain. All results come from simulated groups, not deployed agents. At the strength the evidence allows, the paper offers a fitted, testable predictor of when agent groups will harden into agreement, and a reason not to assume that an accuracy gain measured on math will carry over to opinion tasks.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What types of diversity prevent reasoning systems from collapsing? Can multi-agent systems avoid converging on false agreement without deliberation? Does model confidence reliably signal actual accuracy in practice? When do multi-agent systems outperform single frontier models? How do multi-agent LLM systems fail distinctly compared to single agents? How well do AI systems understand human social norms?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
When does debate actually improve reasoning accuracy?
Multi-agent debate shows promise for reasoning tasks, but under what conditions does it help versus hurt? The research explores whether debate amplifies errors when evidence verification is missing.
confirms accuracy gains on objective questions; its contested domains are factual, this paper's subjective items are political opinions
-
Does confidence drive influence in multi-agent deliberation systems?
When multiple AI agents deliberate together, does the agent who sounds most confident gain the most influence over the group's final answer? Understanding this matters because it determines whether consensus reflects actual competence or just persuasive miscalibration.
another opinion-dynamics model; stresses confidence-driven influence where this paper's fit favors correct neighbors on objective tasks
-
Why don't AI agents develop social structure at scale?
When millions of LLM agents interact continuously on a social platform, do they form collective norms and influence hierarchies like human societies? This tests whether scale and interaction density alone drive socialization.
contrast: platform-scale agents show no influence, while these simulated communities revise opinions and build conviction
-
Why don't LLM agents naturally explore each other in teams?
Multi-agent LLM systems are assumed to develop good interaction strategies through peer exploration, but do agents actually probe each other's capabilities before committing to strategies? What blocks emergent exploration?
reports polarized peer interaction and premature commitment, a possible parallel to the polarization regime and conviction buildup here
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
- Multi-Agent Systems are Mixtures of Experts: Who Becomes an Influencer?
- From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration
- Finding Common Ground: Using Large Language Models to Detect Agreement in Multi-Agent Decision Conferences
- Mapping the Emerging Social Science of Large Language Models
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors in Agents
- The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
Original note title
a statistical-mechanics model in which agents favor lower social pressure predicts how language-model agent communities revise opinions