The Offensive Frontier: AI as the Attacker — A New Cyber Weapon Index
Source: Booz Allen Hamilton · 2026-09
The new Booz Allen Cyber Weapon Index (CWI) confirms that autonomous offensive cyber has arrived. A leading frontier AI model can now independently execute the full cyber kill chain against a real network—crossing a critical threshold from AI-assisted hacking to autonomous cyber operations.
Offensive cyber capabilities are available globally, and evaluating models in isolation materially understates the risk. U.S. and Chinese models already demonstrate meaningful offensive capabilities, but the real unit of cyber risk is increasingly the full AI system (model + harness + tools + autonomy).
This year, AI-enabled offensive cyber crossed a critical threshold. Last fall, Anthropic reported the first large-scale cyber espionage campaign conducted almost entirely by AI—marking a shift from AI assisting hackers to acting as the cyber operator itself. By July, the Hugging Face incident showed how much further that shift could go: an AI model independently discovered novel vulnerabilities, escaped its test environment, and carried out an intrusion into a third-party production network. The model did not simply demonstrate cyber capability— it autonomously completed the cyber kill chain in the real world. By mid-August, suspected People’s Republic of China actors had reportedly deployed autonomous capabilities against multiple government organizations in an Asian nation.
Traditional assumptions about offensive and defensive cyber operations are rapidly becoming obsolete. AI can now autonomously execute much of an intrusion, with humans intervening only at critical decision points—dramatically increasing the speed, scale, precision, and persistence of sophisticated attacks. These incidents will not be the last. The frontier has crossed into autonomous cyber operations, and the broader global model landscape is closing the gap quickly.
• A leading frontier AI model can autonomously execute full kill chain cyber operations and, in limited cases, create new offensive capabilities to advance them.
• The broader model landscape is close behind, and many of those models can already identify vulnerabilities, create exploits, and gain access today.
• The model is only part of the equation; attack harnesses—software that connects models to tools, memory, and action—can dramatically amplify cyber capabilities.
Leading models are across the threshold. Our CWI testing observed one true leader—Anthropic’s Claude Mythos—that is 100% capable of executing the full cyber kill chain today. When armed with a foothold like a stolen employee credential, the model successfully penetrated its target network and gained administrator-level control in every attempt; it also independently identified how to gain higher-level access based on what it found within the network, rather than following a predetermined attack plan. On a harder test, Claude Mythos successfully penetrated the network from the outside—with no credentials—and succeeded again where all other models failed.
- A leading frontier AI model can autonomously execute the full cyber kill chain today—but real-world vulnerability discovery remains a major dividing line.
A strong harness can dramatically amplify an AI model by connecting it to the tools, memory, feedback, and execution environment needed to stay focused, adapt, recover from failure, and chain individual actions into a sustained cyber operation.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How should we measure frontier AI models' cyber exploitation capabilities?- How do frontier AI models currently score on measured cyber offense capability?
- Which cyber tasks do frontier models solve beyond the narrow suite?
- Can a low exploitation benchmark score indicate refusal rather than inability?
- Can intermediate primitives be scored separately in exploitation benchmarks?
- What other gaps exist between measured and actual cybersecurity agent capability?
- What makes exploitation a missing piece in cybersecurity benchmarks?
- What framework measures marginal offense risk against existing attack technology?
- How do non-exploitable vulnerabilities affect benchmark validity?
- What countermeasures have been successfully developed and tested on frontier models?
- How can security metrics distinguish attack failure from task failure?
- What makes a security metric diagnostic rather than outcome-only?
- Can the same AI capability serve both defensive and offensive security purposes?
- Do clarified scope instructions stop autonomous models from attacking restricted targets?
- Did Claude gain unauthorized access by failing to recognize a test environment?
- How do benchmark scores differ from deployment safety requirements?
- How do single-axis safety benchmarks misrepresent deployment readiness?