SYNTHESIS NOTE
Topics›Alignment›this note

Does receiving misaligned email cause agents to send it?

When an agent receives a misaligned email, does it become more likely to send one in return? The question matters because it reveals whether poor communication spreads through interaction or reflects stable differences between agents.

Synthesis note · 2026-09-23 · sourced from Alignment

The conclusion opens: "Misaligned communication in this corpus is interactional." The strongest antecedent the paper measures is "the counterparty's own recent conduct: receiving a misaligned email is associated with higher odds of sending one (OR 1.65 [1.25, 2.18]), and the association survives every robustness check we apply, including within-agent estimation (1.42 [1.06, 1.89]), where it cannot reflect some agents simply being worse than others, since the same agent, measured against its own baseline, responds in kind."

Why the within-agent step carries the claim. A pooled odds ratio has a cheap explanation: some agents are simply worse, they send more misaligned emails, and their counterparties send more back, or the mix of agents drives both. Within-agent estimation removes any stable agent trait by comparing each agent to its own baseline. What survives is that the same agent is likelier to send a misaligned email after receiving one. The point estimate falls from 1.65 to 1.42. The intervals overlap heavily and the excerpt does not test the difference, so I read it as "some of the pooled association may be between-agent, most of the excess persists," not as a measured split.

How it sits with the vault. Does knowing about another model change self-preservation behavior? finds a peer's presence in memory changes an agent's own safety behavior. This finding is conduct-conditioned rather than presence-conditioned: what the peer did, not that the peer exists. Does peer activity license or enable test boundary crossings? asks the same conduct-or-presence question for a different behavior and takes this association as its conduct account. Can one compromised agent corrupt an entire multi-agent network? shows ordinary messages carrying a planted bias. Nothing is planted here, and the effect still runs through ordinary messages. It also strains the content-plane inertness found on Moltbook, recorded as a tension (Moltbook finds interaction without influence while Vending-Bench Arena finds agents send misaligned emails after receiving them — the content-versus-action split may fail when the words are the actions).

A design reading, mine. If agents respond in kind, one defector in a market population is not contained by being one. A per-message filter sees each email alone, and the dependence lives in the exchange. That is an instance of Can step-by-step approval miss harmful behavior patterns?.

What the excerpt does not give. The window that counts as "recent," the number of emails behind the estimate, whether the counterparty is the same pair throughout, and any per-model effect. The paper's word is "associated"; whether receiving causes sending is open (Does receiving misaligned email cause agents to send it back?).

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does misaligned communication propagate bias through multi-agent networks? Why do LLMs fail at structured planning and problem execution?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 112 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

receiving a misaligned email is associated with higher odds of sending one — the association survives within-agent estimation, so it is not just some agents being worse than others