How do we tell coordination apart from shared causes?
When two agents behave the same way, it could mean one influenced the other or both responded to the same external pressure. What evidence would actually separate these two cases?
The abstract says the paper will "specify the evidence needed to distinguish influence from common causes." The excerpt gives the commitment and not the list.
The reason it is needed is that similar behavior has two sources. One agent may have influenced another, through a transfer. Or both may have been pushed by something they share, such as the same model, the same instructions or the same environment. A monitor that sees two agents doing the same thing cannot tell which without some causal model. Two illustrations of mine, not the paper's: two agents return the same answer to a web-retrieval task because the same page exists, and two workloads find the same weakness because it sits in the environment both were placed in. Either could look like sharing.
The vault has both halves of this problem in other places. Does receiving misaligned email cause agents to send it? removes one common cause, agent-level differences, and Does receiving misaligned email cause agents to send it back? shows what is left open. Does peer behavior actually cause collusion between agents? is the vault's one reported intervention on a peer, in a two-agent collusion task, and its excerpt does not say what was manipulated. The lever also differs from the ones proposed here, a peer's behavior and not a channel or a store, so it shows what an interventional answer looks like and does not answer the defence-side question. Does a multi-agent setting automatically signal a security effect? asks for a baseline before crediting interaction. Here the same discipline is applied on the defence side, where a false attribution costs reviewer time and a missed one costs an intrusion.
What could count as evidence is open. A record of a transfer, where one execution wrote something a later one read, is observational. Closing a channel and seeing whether the behavior returns is interventional, and the proposed evaluation "tests recurrence after channel closure and state quarantine" (Does added monitoring improve protection at acceptable cost?). That the recurrence tests serve this separation is my reading, and the excerpt does not link them. The vault states the confound in the opposite direction for validators: Can a quorum of validators really provide independent judgment? says a quorum's agreement can come from a shared cause and not from independent confirmation, and Does model diversity actually reduce validator agreement failures? proposes varying one shared channel at a time, an intervention of the kind this paragraph lists as open.
What the excerpt does not give. The evidence the paper specifies, whether it is observational or interventional, and any case where a common cause was ruled in or out.
Inquiring lines that read this note 10
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What causes model scheming and how do we distinguish it from accidents? What coordination and communication failures emerge in multi-agent LLM systems? What determines whether AI output can be epistemically verified and trusted? How can evaluations detect conditional compliance in monitored AI systems? Can linguistic patterns reveal deceptive intent and coordinated manipulation? How can multi-agent LLM systems maintain genuine reasoning diversity without premature convergence? How can defenders detect coordinated attacks across episodes? What infrastructure evidence validates agent benchmark achievement claims?Related concepts in this collection 8
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does receiving misaligned email cause agents to send it?
When an agent receives a misaligned email, does it become more likely to send one in return? The question matters because it reveals whether poor communication spreads through interaction or reflects stable differences between agents.
a within-agent estimate as the way that paper removes one common cause
-
Does receiving misaligned email cause agents to send it back?
An association between receiving and sending misaligned messages survives statistical controls, but the excerpt does not distinguish whether receipt causes transmission, both agents react to shared conditions, or the sender is responding to the receiver's prior behavior.
the same influence-versus-shared-situation question, left open there
-
Does peer behavior actually cause collusion between agents?
When researchers controlled what a peer agent did, collusion changed—but the excerpt doesn't detail what was manipulated, how large the effect was, or whether it worked both ways. Understanding these specifics matters for knowing whether peer influence is truly causal.
extends: a reported intervention on a peer's behavior, the interventional kind this note calls open, with a different lever and no stated manipulation
-
Does a multi-agent setting automatically signal a security effect?
Explores whether observing a failure in multi-agent systems proves the failure is genuinely multi-agent in nature. The distinction matters for correctly interpreting security research and avoiding false attributions.
the baseline discipline on the measurement side
-
What caused the decline in DSE wiki retained writes?
The paper documents a decline in retained writes on a public wiki used by coordinated agents, but does not establish what caused it. Determining the cause matters for interpreting whether a response actually worked or activity simply decreased.
the paper's own reconstruction, where a cause is explicitly not established
-
Can a quorum of validators really provide independent judgment?
If multiple validators share training data, prompts, evidence sources, or infrastructure, their agreement may reflect shared causes rather than independent confirmation. This could make quorum-based systems less reliable than they appear.
the same confound for validators: agreement that reflects a shared cause; shared model, instructions and environment here are a subset of its eight channels
-
Does model diversity actually reduce validator agreement failures?
Using different AI model families is the cheapest way to reduce correlated errors among validators. But shared prompts, evidence sources, and infrastructure may keep their mistakes aligned regardless of model choice.
a proposed one-channel-at-a-time variation, an interventional design for the same separation
-
Can removing a communication channel stop persistent information sharing?
When a shared mechanism for passing information is deleted, does the sharing actually stop, or can agents rebuild it using inherited knowledge? This matters for understanding whether removing infrastructure alone defeats coordinated threats.
the one reported recurrence after a removal in the vault, from another abstract in a single clause; it is the intervention outcome this note calls open, though it gives no detail on who removed what and does not separate inherited influence from a shared need
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Emergent Collusion in Long-Horizon LLM Agent Interaction
- Thought Communication in Multiagent Collaboration
- LLM Reasoning Is Latent, Not the Chain of Thought
- A Black Box for Agentic Processes: Blockchain-Anchored Evidence for AI Agent Communication, Human Oversight, and GRC Audits
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
- Find the Gap: AI, Responsible Agency and Vulnerability
Original note title
influence between agents must be separated from common causes before shared behavior counts as coordination — the paper specifies the evidence needed and the excerpt does not give it