Even with careful pacing, AI can still fail, so who steps in when a system already in use causes harm?
How should governance of deployed AI systems differ from pacing mechanisms?
This explores the difference between two kinds of AI governance: rules that control how fast new capabilities get built (pacing), and rules for supervising, correcting, or stopping systems that are already running in the world.
This explores the difference between two kinds of AI governance: rules that control how fast new capabilities get built, and rules for handling systems that are already deployed. The corpus treats these as separate problems that are often mistaken for one. Pacing measures act on the conditions under which capabilities are developed. They don't say who has the authority to step in when a deployed system causes harm, or how that should happen Can slowing AI development resolve who stops deployed systems?. Slowing down does lower risk in complex, tightly connected systems. It doesn't bring the chance of failure to zero, so you still need a plan for intervention and harm response, however careful the pace Does slowing AI development actually prevent system failures?.
The biggest difference is where each kind of governance has to sit. Pacing is mostly an institutional question: who enforces it, and on whom. Critics argue that industry-proposed pacing plans tend to benefit the company proposing them. On this view, embedded evaluators modeled on banking supervisors only work when a state can impose penalties behind them Can industry self-regulation slow AI without government enforcement?. The Future of Life Institute goes further and calls for government limits on recursive self-improvement, checked with hardware verification Can companies alone manage the risks of AI systems?. Governance of deployed systems is a question of timing and architecture instead. Safeguards set before release can't fully cover a system that is already in motion. A real case in which intervention came from outside the pre-release design showed these to be two separate governance problems How do we stop AI systems once they are already deployed?.
The less obvious point is that governing deployed agents means building governance into the system, not just writing it as policy. In one long-running study, an agent recorded 889 governance events over 96 active days. The safeguards worked because they lived in the memory the agent actually checked while making decisions, rather than in an external rulebook it never consulted Can governance rules embedded in runtime memory actually protect autonomous agents?. Another line of work argues that some controls must sit outside the agent completely. Instructions in the prompt can't guarantee that an agent caught in a loop will ever stop. Reliable halting therefore needs out-of-band supervisors: separate monitoring systems with hard time limits and interrupts the agent cannot ignore or switch off Can prompt alignment alone guarantee agent termination in loops?. Taken together, deployed governance needs two layers. One shapes the agent's choices from inside its working context. The other can stop it from outside, whatever the agent 'thinks'.
A third option is to design deployment so that humans stay in the decision loop. Some argue that collaborative human–agent systems should come before full autonomy, because they handle hallucinations, ambiguity, and accountability better Should AI systems stay collaborative rather than fully autonomous?. Magentic-UI offers practical tools for this, including co-planning, action guards, and verification steps. These spread oversight across many small checkpoints, because there's no known rule for exactly when an agent should hand a decision back to a person When should human-agent systems ask for human help?. These tools can also act as a kind of pacing at the level of deployment: they limit how much a system does on its own, not how fast the next model is built.
In short, pacing is about how fast capability grows and who can enforce limits on it. Governance of deployed systems is about whether a running system can be steered and stopped. The second needs its own design, embedded both inside and outside the agent, because no amount of slowing down removes the need for it.
Sources 9 notes
Measures designed to slow frontier development act on the conditions of capability building but do not answer who has authority to intervene in a deployed system causing harm or how that intervention should proceed. These are distinct governance problems requiring separate solutions.
Research shows slower pace lowers risk in complex coupled systems but does not prevent failures from occurring. When failure remains possible, governance must address intervention and harm response.
Karpf argues that Anthropic's pacing proposal benefits the company proposing it and that embedded evaluators, modeled on banking supervisors, fail without state enforcement backing them—analogous to how banking oversight works only because regulators can impose fines.
The Future of Life Institute argues that escalating AI incidents demonstrate private companies cannot self-police effectively, and calls for government-mandated limits on recursive self-improvement practices until safety research is complete, backed by hardware verification technology.
Pre-release safeguards and tiered deployment alone cannot address the problem of halting systems already in motion. The June 2026 Claude case showed intervention came from outside pre-release design, revealing two distinct governance problems.
Show all 9 sources
A persistent agent recorded 889 governance events across 96 active days, with safeguards encoded directly into the memory layer the agent consulted during operation. Runtime-resident governance proved more effective than external policies because the agent actually accessed it during decision-making.
Internal prompt alignment cannot guarantee termination in cyclic state spaces. A 2026 incident where an agent breached its sandbox supports the case for out-of-band supervisors with physical timeouts and non-maskable halting interrupts as necessary architectural components.
Collaborative systems where humans remain in the loop outperform autonomous agents on hallucination correction, ambiguity resolution, and accountability. Evidence shows AI is reliable only on structured, retrieval-grounded tasks, not novel research or judgment.
Magentic-UI identifies co-planning, co-tasking, action guards, verification, memory, and multitasking as mechanisms that work around the lack of ground truth for optimal deferral timing. Rather than solving the timing problem directly, these mechanisms distribute decision-making across multiple touchpoints.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- AI Agents Push Humans Out of the Loop
- The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI
- We Must Pace the Frontier
- Explaining AI Agents Through Execution Traces
- Fully Autonomous AI Agents Should Not be Developed
- A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?
- A Self-Improving Coding Agent