SYNTHESIS NOTE
Topics›Agents Multi›this note

Do teams of personal agents outperform a single coordinator?

When multiple users each have their own AI agent in a shared environment, do they achieve better outcomes than having one agent serve everyone? This matters for understanding how to scale agent delegation across groups.

Synthesis note · 2026-10-08 · sourced from Agents Multi

"Worse Together" studies what happens when different users each delegate to their own agent inside a shared environment — "an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff." Across five frontier models and 77 scenarios, it compares a coordinator (one agent serving every user) against a silent team (one agent per user, no communication) and a peer-to-peer team (one agent per user, with communication). "Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps." In the personal assistant environment, "the coordinator fulfills a targeted user request about twice as often as teams."

The paper attributes the gap to behaviors that arise from multi-user contention rather than task difficulty — environments were chosen so "each individual request is within the capabilities of current frontier models, so that differences between formations reflect coordination rather than task competence." It names three failure behaviors: "stalling as teams grow, overriding each other's actions, and fabricating claims." Mitigations are environment-specific rather than general: "a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing" recover some performance, but the paper stresses that "one cannot assume either" a team lead or added guidance "exists" in shared workspaces that arise on their own.

This sits alongside When do multi-agent systems actually outperform single agents? but for a different reason: that paper's node-, edge-, and path-level defects degrade a multi-agent system on one shared task as single-agent capability rises; this paper holds task difficulty constant and instead varies who the agents represent, finding the coordinator beats the team because the agents serve different users with uncoordinated goals, not because any one agent is individually weak. Its mitigation finding — procedural instructions and a read-before-commit platform check outperform ad hoc peer messaging — points the same direction as Does structured artifact sharing outperform conversational coordination?: structured protocol beats free-form agent-to-agent dialogue. It also extends Where do user values break down in agent supervision? by giving a second route to the same supervision gap: a run can go wrong not only because one user cannot see what their own agent is doing, but because a second user's agent, with no visibility into the first user's goal, can override or stall it through shared-resource contention alone.

The results come from four constructed benchmark environments (to be released as MAMUBench, 74 scenarios), scored mechanically from environment logs, with consent and fabrication labeled by Claude judges the authors say they "did not validate against human judgments." The paper itself calls running all formations "expensive, long, and stochastic" — one Opus 5 peer-to-peer pass through MAMUBench runs about 3.4 billion tokens — and notes "no standard harness yet exists for multi-user, multi-agent systems." So this is a finding about five tested models across four designed scenarios, not a measured failure rate in deployed systems. The paper's own extension to deployment — shared Slurm clusters, merge queues, DoorDash group carts — is an inference that the same structural pressure, competing users on one resource, already exists there, offered as a reason to study it rather than evidence that it is failing there now.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

When do multi-agent systems improve over single frontier models? How should humans and AI agents share control and decision-making?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 136 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

multi-user multi-agent teams perform worse than a single coordinator across four shared-resource environments — two collapse without a channel