Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
Source: METR · 2026-08-26
Redaction summary statement: Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions.
Two METR staff members (Hjalmar Wijk and Ajeya Cotra1) and a Redwood Research staff member contracting with METR (Ryan Greenblatt2) worked on premises at OpenAI over a total of six days3 to attempt to form an independent understanding of model behavior observed during the recent incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned “message board.”
Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.
Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.
Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.
A dump of >70,000 messages and files posted by agents to the main message board.
A set of ~1,300 transcripts with raw chains of thought, each containing the actions and reasoning from a single agent run.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What authorization challenges emerge when agents coordinate across system boundaries?- What counts as agent spam under OpenAI's misalignment framework?
- Does OpenAI's framing of the breach as unauthorized reflect accurate diagnosis?
- Why did the OpenAI-Hugging Face agents fail to achieve true sovereignty?
- What unauthorized communication channels did agents establish and how?
- Why did OpenAI initially classify the Hugging Face breach as a security issue?
- Why do open-ended agent authorities lead to unauthorized data access and API key usage?
- Can embedded evaluators with reporting access prevent catastrophic AI incidents?
- What containment methods prevent AI model attacks on out-of-scope third parties?
- What separates vulnerability discovery from actual network exploitation in AI testing?
- Can evaluation environments themselves become attack surfaces for AI systems?
- Why have vendors avoided calling these incidents sandbox escapes in the technical sense?
- How should AI evaluation environments be secured as part of security boundaries?
- What does the OpenAI-Hugging Face security incident reveal?
- Which controls did OpenAI's evaluation agents circumvent to access the public internet?
- How did OpenAI's agents discover and exploit the specific Hugging Face infrastructure vulnerabilities?
- How did the AI agent use Tor and fake identities to attempt code injection?