The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI

Paper · arXiv 2609.22882 · Published September 19, 2026
LLM Alignment

Introduction. The next problem for artificial intelligence (AI) governance is not only how to regulate AI systems before they are released, but how to stop them once they are in motion. The Mythos episode of June 2026 shows what is at stake. Anthropic had built its most capable class of model yet, strong enough at finding and exploiting software vulnerabilities and at reasoning about biology and chemistry to be valuable to cyber-defenders and drug developers and dangerous in the wrong hands.1 Anthropic split the release in two. Claude Mythos 5, with some safeguards lifted, went only to a small group of cyber-defenders and infrastructure providers under a program run with the government;2 Claude Fable 5, released publicly on June 9, 2026, was the same model behind classifiers that diverted requests touching cybersecurity, biology and chemistry to a weaker model. On June 12, 2026, three days after the public release, the government intervened.

Discussion / Conclusion. The Anthropic and OpenAI episodes with which this Article opened illustrate the consequences of this regulatory gap. An export-control directive restricting foreign access caused Anthropic to suspend both models globally. The evidentiary basis for the directive was not disclosed, and access was restored without a stated justification. Hugging Face terminated the intrusion by an OpenAI agent through its own security measures, before the source of the intrusion had been identified. The ensuing debate has focused principally on the pace of development. Dario Amodei’s proposal would slow the frontier through embedded evaluators and coordinated capability checkpoints.243 Such measures govern the conditions under which capabilities advance. They do not resolve who may intervene when a deployed system causes harm, or how that intervention should proceed. Slower development may reduce risk, but it cannot eliminate the possibility of failure in complex, tightly coupled agentic systems.244 The incident record examined in Part III indicates the scale of the deficiency: of 1,213 coded incidents, four in five record no stop of any kind, and where no usable mechanism existed the missing element was more often legal or institutional than technical.

Lines of inquiry this paper opens 3

Research framings built by reading the notes related to this paper — the questions it feeds into.

How should models express uncertainty rather than forced confident answers? How do we evaluate AI systems when user perception misleads actual performance? How can identical external performance mask different internal representations?