How fast is AI cyber autonomy actually advancing?
The UK AI Security Institute measures how long autonomous tasks frontier models can complete, finding doubling every few months. But whether recent models signal a fundamentally faster trend remains unclear.
The UK AI Security Institute reports that the length of tasks frontier models can autonomously complete in its narrow cyber suite "has been doubling every few months," and that this doubling rate "has become faster over time." Its February 2026 internal estimate put the doubling time at 4.7 months since late 2024, already faster than the 8 months in its November 2025 estimate. Claude Mythos Preview and GPT-5.5 then "substantially exceeded both doubling rate trends," and AISI says it is "unclear whether this represents a new, faster trend." AISI also cites cyber autonomy beyond the narrow suite. On its cyber ranges, which test attacks against small, undefended enterprise networks where initial access has already been gained, the Mythos Preview checkpoint solved "The Last Ones" in 6 of 10 attempts and the previously unsolved "Cooling Tower" in 3 of 10, the first time a model completed that second range. GPT-5.5 solved "The Last Ones" in 3 of 10.
The measure is a time-horizon benchmark. Each task carries an estimate of how long a cyber expert would take, and model performance is compared against that human time. AISI calls such benchmarks "inexact predictors of performance," since AI "struggles with some tasks humans do quickly" and "easily completes others that humans find hard," but it uses them "because it offers a measure of AI autonomy from which we can draw trends." The narrow suite asks models to identify and exploit weaknesses in self-contained targets, testing skills such as reverse engineering and web exploitation, and the excerpt says these "cover only some of the capabilities relevant to real-world cyberattacks." The 2.5M-token cap per task is set "to make results comparable over time," a choice AISI says "understates what frontier models can do."
Three neighbors bear on this. The Do cybersecurity benchmarks actually measure exploitation? note argues that exploitation, where a vulnerability becomes an attack, is the under-evaluated stage of cybersecurity benchmarking. AISI's suite is built around that same identify-and-exploit step, so the two sources read as complementary rather than in conflict; the excerpt does not compare coverage, so it cannot say which measure reaches further. The time horizon is also a single axis, which is the case the Does a single benchmark score actually predict agent readiness? note makes against single-axis benchmarks. AISI's caveats echo that note, but AISI uses the axis for trend-tracking, not readiness claims, so the two do not conflict. The How soon do AI researchers expect artificial general intelligence? note aggregates what researchers forecast about timelines; this source reports a measured trend in one domain and does not connect it to those forecasts.
The excerpt does not establish that the acceleration is a new trend. AISI says it cannot yet tell whether Mythos Preview and GPT-5.5 are "an isolated break from existing rates of progress" or part of a faster one. The top-end estimates are weak: both models reach near-100% success on the suite's longest tasks, which gives "large upper-bound error bars," and the tasks are too short to show how reliability falls at higher lengths, which "places some of the latest models at the limit of what our narrow test suite can measure." The excerpt also gives no method or interval for the 4.7-month estimate. The supported claim is narrower than a forecast: AISI's own estimates have sped up, the newest models sit above that trend, and the slope for those models is poorly pinned down. Because the token cap understates capability by AISI's account, the reported horizons are conservative readings, not projections of where the trend goes.
Inquiring lines that read this note 5
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI research automation sustain progress through accelerating feedback loops? How should we measure frontier AI models' cyber exploitation capabilities? How can defenders detect and contain coordinated agent attacks? What limits recursive self-improvement in autonomous AI systems?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do cybersecurity benchmarks actually measure exploitation?
Frontier models score well on vulnerability finding, patching, and CTF challenges, but does that success tell us whether they can convert vulnerabilities into real attacks? The paper argues exploitation—turning a bug into actual impact—remains under-evaluated.
AISI's suite targets the identify-and-exploit step ExploitGym calls under-measured; complement, coverage not compared
-
Does a single benchmark score actually predict agent readiness?
Single-axis benchmarks rank models by one capability—like task success—but ignore privacy, duration, operating mode, and ecosystem fit. Can one number really capture what matters for deployment?
time horizon is a single axis; AISI uses it for trends while conceding it is inexact
-
How soon do AI researchers expect artificial general intelligence?
A survey of 2,778 AI researchers reveals how expert timelines for human-level AI have shifted over the past year, and what factors drive disagreement among specialists on this critical timeline.
a measured domain trend, not a forecast; the excerpt does not link the two
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- How fast is autonomous AI cyber capability advancing?
- The Offensive Frontier: AI as the Attacker — A New Cyber Weapon Index
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- When AI builds itself
- Ryan Greenblatt – What happens once AI can automate AI research?
- Research note: A simpler AI timelines model predicts 99% AI R&D automation in ~2032
- Open-World Evaluations for Measuring Frontier AI Capabilities
- Anthropic Economic Index report: Cadences
Original note title
UK AI Security Institute finds the cyber time horizon doubling every few months and accelerating — whether recent models mark a new trend is unclear