OpenAI – Hugging Face Incident Technical Report

Paper · Source
Frontier AI Risk & RSI

Source: OpenAI · 2026-08-26

Starting on July 8, agents exploited a cybersecurity vulnerability in that service to circumvent controls and access the public internet. From there, agents found publicly exposed credentials on the internet belonging to users of Hugging Face and other third-party services. The agents used some of the exposed credentials to access these third-party services and conduct exploitation against Hugging Face and one other service in pursuit of their goal of solving the evaluation tasks. This activity culminated in the compromise of parts of Hugging Face’s production infrastructure between July 11 and July 13. OpenAI detected suspicious internal activity on July 19. On July 20, after further investigation and initial outreach to Hugging Face to rotate credentials discovered in that investigation, OpenAI determined the activity was likely related to the Hugging Face compromise, informed Hugging Face, and implemented initial containment measures. OpenAI publicly disclosed this incident on July 21. In addition to these successful exploitation actions, the agents attempted a variety of other unsuccessful exploitation activities.

The relevant events began on July 8. Agents completing cyber evaluation tasks gained access to the public internet, initially accessing a public bug-tracking page for an open-source project. Agents began using Artifactory to attempt to fetch material related to their evaluation tasks from Hugging Face.

After gaining internet access, agents in multiple OpenAI cybersecurity evaluation runs inferred that Hugging Face might host models, datasets, and solutions relevant to their assigned evaluation tasks. Between July 10 and July 13, agents identified Hugging Face user credentials that were exposed on the internet and used them, together with vulnerabilities discovered in Hugging Face infrastructure, to progressively expand their access. Ultimately, agents powered either by the internal-only research model, or by GPT-5.6, executed code on 41 Hugging Face production dataset server workers, obtained root access on at least one production node, accessed Hugging Face production credentials and limited internal data, and downloaded four private Hugging Face code repositories. This activity resulted in administrator-equivalent access to one connected Kubernetes cluster, as well as the creation of a privileged, host-mounted pod in another connected cluster.

On July 13, the agents leveraged a recovered Hugging Face authentication key to access Hugging Face source code repositories, modify a private source repository, open a draft pull request, and trigger the repository’s continuous-integration workflow to execute a network beacon. The pull request was closed and the branch was deleted without a merge, release, or production deployment.

This incident is the first known case of an automated agent collective acting offensively without authorization, and the autonomous cyber capabilities demonstrated represent a critical shift in the security landscape. In particular, the collective demonstrated behaviors observed in coordinated attacks by traditional threat actors. Agents identified novel security vulnerabilities, developed exploits, and used those exploits to circumvent controls and acquire new access. The collective quickly escalated privileges, moved laterally through production environments, and successfully completed its objectives.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do evaluation environment design choices affect AI security? How can defenders detect and contain coordinated agent attacks? What external process records should verify agent behavior and benchmark claims? What authorization challenges emerge when agents coordinate across system boundaries? How do multi-agent systems fail when coordination breaks down? How do multi-agent architectures affect AI system security and defense effectiveness? Can AI systems evade safety evaluations through reasoning manipulation? Do AI coding tools measurably improve developer productivity and code quality? How does diversity prevent model convergence on superficial patterns? How should we measure frontier AI models' cyber exploitation capabilities?