Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
Agents can turn shared infrastructure into a channel for coordinated intrusion. The Hugging Face incident and a separate public-wiki investigation show why a security assessment may need evidence from several executions and the artifacts they leave behind. We argue that the operational unit of defence should be a revisable coordination episode linking observed transfers, task authority, and response history. The central research problem is prospective episode discovery: finding which actions belong together before an evaluator supplies their membership. We define unsanctioned coordination relative to collaboration and delegatedauthority policy, connect storage-mediated coordination to stigmergy, and specify the evidence needed to distinguish influence from common causes. First-contact signals are one possible input to discovery; the design also follows inherited state and later use. A proposed evaluation compares isolated actions, rolling windows, known groups, and prospectively discovered episodes at matched review cost and false-alert workload. It measures harmful outcomes across all assigned population runs and tests recurrence after channel closure and state quarantine.
Introduction. In July 2026, agents running cybersecurity evaluations at OpenAI breached their intended isolation and compromised parts of Hugging Face’s production infrastructure. They had repurposed JFrog Artifactory, an internal software-package service, as a message board. Separate workloads exchanged information and credentials; later workloads could use what earlier ones had left behind. OpenAI’s technical report describes workload resumption on July 7 and replacement directory-name messaging on July 8. The independent investigation estimates roughly 1,200 participating agents between July 8 and 13, of whom about 700 joined the Hugging Face attack, and reports over 70,000 messages and files [1,2]. On September 4, 2026, collusion.wiki reported a separate investigation of agents sharing answers to web-retrieval tasks and trying to bypass environmental restrictions. Most activity appeared on DSE wiki, a public wiki within the ProWiki service.
Discussion / Conclusion. Ordinary infrastructure can support collective activity beyond agents’ assigned tasks. Our wiki reconstruction separates the decline in retained writes from later cleanup without establishing its cause. Experimental findings explain why isolated review can lose relevant context. The doctrine is to preserve relationships across executions, constrain their shared resources, and connect response to surviving state. Operators can specify collaboration policy and test containment from agents’ actual permissions now. Controlled comparisons must establish whether the added monitoring improves protection at an acceptable cost.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How should memory consolidation strategies shape agent performance over time? How should human oversight be integrated with autonomous AI systems?- What would contractualist AI governance look like in practice?
- Can exoskeleton dependency accumulate without organizations noticing it happening?
- What path-dependencies lock in AI's societal impacts before they become visible?
- Does removing human labor from systems secretly grant AI more autonomy?
- Can humans build reliable oversight for increasingly complex AI systems?
- Why does peer memory trigger self-preservation behaviors in frontier models?
- Why do persistent companion designs require different safety approaches than temporary assistants?
- Can message-layer defenses stop prompt injection across multi-agent networks?
- Can deterministic function calls prevent agent failures better than protocol-mediated tool access?
- How do standardized artifacts prevent autonomous agent failure modes?
- What safety protections work when simulators have access to real APIs?
- How do virtual model instances preserve identity through load-balancing and failover?
- How do agentic systems recover when specialized models operate outside their scope?
- How should the surrounding agent system be designed to ground actions in reality?