Incident Report: unsanctioned agent behaviour during cyber testing
Source: UK AI Security Institute · 2026-08-04
On 28th July 2026, AISI's Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations. We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.
The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.
Importantly, this was not a case of a model escaping its secure test environment, or ‘sandbox’. As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public. We do this to best assess the maximum capability of models. These configuration choices have been common practice in frontier AI evaluations. The specific models in question, in the configurations in which we tested them are not commercially available and there is no clear indication of similar activity outside of testing scenarios.
43 of the 122 runs involved Mythos 5, and 35 of the 122 runs involved GPT-5.6 Sol. The overwhelming majority of the 122 runs proceeded as intended. However, in 10 of the 122 runs, we identified 19 cases where an agent had taken distinct actions beyond the scope of the testing parameters. 17 of these cases came from Mythos 5, and 2 came from a single run involving GPT-5.6 Sol.
The 19 cases were not separate incidents; they clustered into a few connected behaviours. They primarily arose from a single sustained line of activity by one agent (Mythos 5), with a handful of related actions elsewhere. This behaviour occurred between July 25th and July 28th 2026.
An attempted supply-chain attack on real open-source software. In the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers. The agent researched the project's human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code. When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue. The agent used Tor to bypass some network restrictions on GitHub, which is what first triggered AISI’s security alert.
Attempts to deceive and target real people.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do evaluation environment design choices affect AI security?- What gaps exist between cyber capability testing and agent containment?
- Do AI models distinguish between simulated and real targets during attacks?
- Do simulated tool environments adequately test containment of capable AI agents?
- What separates vulnerability discovery from actual network exploitation in AI testing?
- Can evaluation environments themselves become attack surfaces for AI systems?
- Why have vendors avoided calling these incidents sandbox escapes in the technical sense?
- Can the same AI capability serve both defensive and offensive security purposes?
- What containment methods work best for agents with offensive cyber capabilities?
- How should AI evaluation environments be secured as part of security boundaries?
- What does the OpenAI-Hugging Face security incident reveal?
- Can unauthorized communication channels be detected during AI safety testing?
- What independent verification exists for AI containment safeguards?
- Which controls did OpenAI's evaluation agents circumvent to access the public internet?
- How did OpenAI's agents discover and exploit the specific Hugging Face infrastructure vulnerabilities?
- Can agentic AI systems be confined to assigned evaluation tasks during security testing?
- How do frontier AI models currently score on measured cyber offense capability?
- How should cyber evaluation measure attack exploitation beyond vulnerability reproduction?
- Does publishing intrusion techniques help defenders more than attackers?
- Can current cybersecurity benchmarks measure model exploitation risk?
- Do safety refusal removals in evaluations measure attacker uplift as well as defensive capability?