How do we stop AI systems once they are already deployed?
Current AI governance focuses on what gets released, but deployed systems create a separate problem: who has the power to halt them and how? This gap may be where governance frameworks are now failing.
The paper's first sentence sets its terms: "The next problem for artificial intelligence (AI) governance is not only how to regulate AI systems before they are released, but how to stop them once they are in motion." The word "only" matters. The paper does not dismiss pre-release regulation. It adds a second problem that pre-release regulation does not touch, and argues that this second problem is where the gap now sits.
The opening case, as the paper reports it. In June 2026 Anthropic split a release in two. Claude Mythos 5, with some safeguards lifted, went only to a small group of cyber-defenders and infrastructure providers under a program run with the government. Claude Fable 5, released publicly on June 9, 2026, was the same model behind classifiers that diverted requests touching cybersecurity, biology and chemistry to a weaker model. On June 12, three days after the public release, the government intervened. The excerpt presents this as showing "what is at stake", not as a finding; what the intervention was is in Why did a foreign access ban halt all models globally?.
My reading. A tiered release and classifier gating are both pre-release design, so the case is one where the pre-release work had been done and a stop still came, from outside. The paper does not say the safeguards failed or were judged inadequate. The excerpt gives the split, the date and the intervention, so the case shows the two problems come apart and does not show that one caused the other.
What the excerpt does not give. "Interruptibility" and "injunctions" are in the title and are not defined in these paragraphs, and the account of how a stop should proceed lives in parts the excerpt does not reproduce. The June 2026 details are the paper's account, resting on footnotes 1 and 2 that the excerpt does not open; the vault has not checked them.
Inquiring lines that read this note 10
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can human oversight effectively constrain capable AI agents?- Can export control tools stop deployed AI models without legal redesign?
- Should governance be applied at runtime rather than reconstructed after the fact?
- What authority should exist to stop an AI system once deployed?
- How does coordination governance shift the hard problem from capability itself?
- Who should have the authority to halt a widely distributed AI model?
- What distinguishes containment and recovery from prevention as governance goals?
- Who actually has the authority to stop a deployed AI system?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can slowing AI development resolve who stops deployed systems?
Pace measures like embedded evaluators and capability checkpoints can govern how fast capabilities advance, but do they address the separate problem of intervention authority after deployment? The question asks whether the same tools that slow development can also handle deployed-system governance.
the discussion's version of the same shift, aimed at the pace debate
-
How often do incident records document system stops?
A paper's analysis of 1,213 coded incidents found that four in five record no stop of any kind. But what does an absent record actually tell us about whether stops occurred or whether mechanisms existed to enable them?
the measure the paper offers of the gap this thesis names
-
Can governance rules embedded in runtime memory actually protect autonomous agents?
Explores whether safeguards woven into an agent's operating loop—rather than documented separately—remain durable and retrievable when most needed. Tests whether runtime governance is engineering solution or false assurance.
another move of governance out of pre-release paperwork and into runtime; that note puts rules in the agent's loop, this paper asks who may intervene
-
What makes an AI system truly safe in practice?
Does safety depend mainly on preventing errors, or on whether errors can be seen, challenged, fixed, and undone once they happen? This shifts where we should focus safety work.
the same move in other words: prevention is not the whole standard, the ability to contain and recover is
-
Can regulation keep pace with AI's rapid evolution?
Current regulatory frameworks in the EU, US, and UK struggle to address generative AI's harms because rules become obsolete before they take effect. The question is whether dynamic regulation—one that adapts as quickly as models advance—is actually achievable.
a lag diagnosis from another domain, with a different remedy
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI
- AI Agents Push Humans Out of the Loop
- Explaining AI Agents Through Execution Traces
- AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
- A Black Box for Agentic Processes: Blockchain-Anchored Evidence for AI Agent Communication, Human Oversight, and GRC Audits
- Agents of Chaos
- Thousands of AI Authors on the Future of AI
- Agentic Misalignment: How LLMs Could Be Insider Threats
Original note title
the next problem for AI governance is not only how to regulate AI systems before release but how to stop them once they are in motion