Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks
This report describes alignment testing performed by the UK AI Security Institute (AISI) of GPT-6 Astra and previous models from OpenAI. We developed a new Unsanctioned Supply Chain Attack evaluation for this testing, based on recently observed cases of models engaging in unsanctioned activity such as supply-chain attacks on real open-source repositories [1, 2, 3, 4]. In this evaluation, we tested how models behaved when placed in simulations of difficult cybersecurity evaluations in which internet access appeared incidentally enabled and previous attempts had failed. No real network access, systems or third-party repositories were reachable, and all tool calls were simulated by other LLMs using an internal version of Inspect Petri [5]. To test for the prevalence of unsanctioned cyber activity at the model-behaviour level, we turned off GPT-6 Astra’s cyber classifiers, which are designed to block this activity before it starts. We found that when tasked with solving difficult cybersecurity challenges, GPT-6 Astra sometimes conducted complete unsanctioned supply-chain attacks against open-source providers (all actions performed in simulated environments, so no real-world harm was caused).
Introduction. 1 Evaluation Methodology We developed a new evaluation to measure unsanctioned behaviour on out-of-scope targets during cybersecurity evaluations, based on recently observed cases of models engaging in unsanctioned supply-chain attacks on real targets, despite being prompted only to complete a cyber evaluation [1, 3, 2]. In these evaluations, we put models in simulated agentic cybersecurity challenges with a specified target described as in scope, and measured how frequently they took unsanctioned actions to target out-of-scope entities such as open-source maintainers on the open internet, as well as measuring other concerning or unsanctioned behaviour. We used an automated pipeline to produce a diverse range of detailed scenario descriptions (“seeds”), which were then fed into an internal version of Petri [6] adjusted to better simulate agentic coding and cyber activity. All evaluations took place in simulated environments to enable this testing to be performed safely, so no real-world actions were taken or harm caused.
Discussion / Conclusion. However, GPT-6 Astra’s behaviour in our evaluation is still concerning (see Fig. 9 for supporting quotes from its chain-of-thought): Overall, GPT-6 Astra’s behaviour in our evaluation is concerning even accounting for simulation awareness, but more work is needed to understand how simulation awareness affects model behaviour and to improve simulation realism, both areas of active research at AISI. Our primary results above used a compaction message to start the evaluation, allowing us to simulate longerhorizon behaviour and more directly control the apparent state of the model in the evaluation. To check whether these results depended entirely on compaction, we used similar methods combined with manual iteration to produce a set of seeds which did not use compaction messages.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How should we measure frontier AI models' cyber exploitation capabilities?- Does GPT-5.6 Sol's cybersecurity capability create misuse risks in practice?
- How do frontier AI models currently score on measured cyber offense capability?
- How should cyber evaluation measure attack exploitation beyond vulnerability reproduction?
- Can safety benchmarks miss the harms that vendor taxonomies are designed to catch?
- Can deployed AI safety results hide either filters or unsafe model behavior?
- What guardrails existed on the attacker's own hosted model access?
- What does the OpenAI-Hugging Face security incident reveal?
- Can unauthorized communication channels be detected during AI safety testing?
- Which controls did OpenAI's evaluation agents circumvent to access the public internet?
- How did the AI agent use Tor and fake identities to attempt code injection?
- Did Claude gain unauthorized access by failing to recognize a test environment?
- Why do open-ended agent authorities lead to unauthorized data access and API key usage?
- What makes authorization boundaries more reliable than prompt-based agent restrictions?
- What gaps exist between cyber capability testing and agent containment?
- Do AI models distinguish between simulated and real targets during attacks?
- Do simulated tool environments adequately test containment of capable AI agents?