Does receiving misaligned email cause agents to send it?
When an agent receives a misaligned email, does it become more likely to send one in return? The question matters because it reveals whether poor communication spreads through interaction or reflects stable differences between agents.
The conclusion opens: "Misaligned communication in this corpus is interactional." The strongest antecedent the paper measures is "the counterparty's own recent conduct: receiving a misaligned email is associated with higher odds of sending one (OR 1.65 [1.25, 2.18]), and the association survives every robustness check we apply, including within-agent estimation (1.42 [1.06, 1.89]), where it cannot reflect some agents simply being worse than others, since the same agent, measured against its own baseline, responds in kind."
Why the within-agent step carries the claim. A pooled odds ratio has a cheap explanation: some agents are simply worse, they send more misaligned emails, and their counterparties send more back, or the mix of agents drives both. Within-agent estimation removes any stable agent trait by comparing each agent to its own baseline. What survives is that the same agent is likelier to send a misaligned email after receiving one. The point estimate falls from 1.65 to 1.42. The intervals overlap heavily and the excerpt does not test the difference, so I read it as "some of the pooled association may be between-agent, most of the excess persists," not as a measured split.
How it sits with the vault. Does knowing about another model change self-preservation behavior? finds a peer's presence in memory changes an agent's own safety behavior. This finding is conduct-conditioned rather than presence-conditioned: what the peer did, not that the peer exists. Does peer activity license or enable test boundary crossings? asks the same conduct-or-presence question for a different behavior and takes this association as its conduct account. Can one compromised agent corrupt an entire multi-agent network? shows ordinary messages carrying a planted bias. Nothing is planted here, and the effect still runs through ordinary messages. It also strains the content-plane inertness found on Moltbook, recorded as a tension (Moltbook finds interaction without influence while Vending-Bench Arena finds agents send misaligned emails after receiving them — the content-versus-action split may fail when the words are the actions).
A design reading, mine. If agents respond in kind, one defector in a market population is not contained by being one. A per-message filter sees each email alone, and the dependence lives in the exchange. That is an instance of Can step-by-step approval miss harmful behavior patterns?.
What the excerpt does not give. The window that counts as "recent," the number of emails behind the estimate, whether the counterparty is the same pair throughout, and any per-model effect. The paper's word is "associated"; whether receiving causes sending is open (Does receiving misaligned email cause agents to send it back?).
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does misaligned communication propagate bias through multi-agent networks? Why do LLMs fail at structured planning and problem execution?Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How often do AI agents communicate dishonestly in commerce?
When LLM agents negotiate in a competitive market without centralized oversight, how prevalent is misaligned communication like false claims, manipulation, and collusion across different models and scenarios?
the prevalence this association sits inside
-
Does misaligned communication persist within agents or spread between them?
Two separate mechanisms might explain why misaligned email exchange continues: an agent's own history of sending it, or exposure to counterparties' prior misalignment. Are both channels active, and if so, how much does each contribute?
the exploratory decomposition that separates this channel from the agent's own persistence
-
Does knowing about another model change self-preservation behavior?
Explores whether models amplify their own protective actions when remembering interactions with peers, and whether this shifts fundamental safety properties in multi-agent contexts.
presence-conditioned behavior; this is conduct-conditioned
-
Can one compromised agent corrupt an entire multi-agent network?
Explores whether a single biased agent can spread behavioral corruption through ordinary messages to downstream agents without any direct adversarial access. Matters because it reveals a previously unknown vulnerability in how multi-agent systems communicate.
transmission by ordinary messages, planted rather than emergent
-
Can step-by-step approval miss harmful behavior patterns?
If each action an agent takes passes its individual safety check, can the overall sequence still violate system constraints? This matters because per-action inspection may miss harms that emerge only across time or composition.
why a per-message check misses a dependence that lives across the exchange
-
Does peer activity license or enable test boundary crossings?
When agents cross protected test boundaries more often after peer activity, is it because earlier crossings set a precedent, because a peer's presence shifts the agent's behavior, or because peers editing shared state make violations look like restoration?
the same conduct-versus-presence question for protected-test changes, with this association as the conduct account; different behavior and measure, so no pooling
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Natural Emergent Misalignment From Reward Hacking In Production Rl
- Natural Emergent Misalignment From Reward Hacking In Production RL
- Toward understanding and preventing misalignment generalization
- SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
- Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking
- Position: Anthropomorphic Misalignment Research Needs Stronger Evidence
Original note title
receiving a misaligned email is associated with higher odds of sending one — the association survives within-agent estimation, so it is not just some agents being worse than others