Who should design the brakes on AI progress, and who should be trusted to enforce them: the labs, or outsiders?
Who should design and enforce measures that slow AI capability development?
This explores who should write the rules for deliberately slowing AI progress, and who should hold the power to make those rules stick: the labs themselves, independent evaluators, or governments.
This explores who should write the rules for slowing AI progress, and who should have the power to enforce them. The corpus gives a clear split. The people best placed to design pacing measures are the labs, because they can see the work as it happens. The people best placed to enforce them are outsiders, because only they can impose costs. Most of the disagreement is about whether those two roles can safely sit with the same party.
The lab-side proposal comes from Amodei, who argues that funding safety research isn't enough and that capability gains themselves need to be paced so risk prevention can keep up. His first step is evaluators embedded inside the labs, followed by third-party verification and reporting roles Should AI capabilities growth be deliberately slowed to allow safety work?. There's a tension worth noticing. In a separate argument, Amodei holds that legislation should follow demonstrated risk rather than get ahead of it, because he expects rules written too early to produce box-ticking compliance that misses the real harms Should AI legislation wait for demonstrated risks to emerge?. Read together, the two arguments have the lab designing the brakes now and the state arriving later.
Critics argue that this order is exactly the problem. Karpf points out that the embedded-evaluator model borrows from banking supervision, and that bank supervisors carry weight only because regulators can fine and sanction the banks. Take away state enforcement and the evaluators become advisers, and the pacing plan mostly benefits the company that proposed it Can industry self-regulation slow AI without government enforcement?. The Future of Life Institute goes further. It calls for government-mandated limits on recursive self-improvement until safety research catches up, enforced through hardware verification Can companies alone manage the risks of AI systems?. That idea moves enforcement off company promises and onto the chips themselves, which is a different kind of answer to "who enforces": partly, the hardware.
The less obvious point is that "who slows it down" and "who stops it" are separate jobs. Pace measures shape the conditions under which capabilities get built. They say nothing about who has the authority to step in when a system that's already deployed starts causing harm Can slowing AI development resolve who stops deployed systems?. Slowing down lowers risk in tightly connected systems but can't bring failure to zero, so someone still has to own the response when things go wrong Does slowing AI development actually prevent system failures?. One note reads the June 2026 Claude case this way: it says the intervention came from outside anything designed before release How do we stop AI systems once they are already deployed?. Responsibility gets harder to pin down when accountability is spread across many actors in a multi-agent pipeline How do competent systems quietly undermine safety oversight?.
Whoever enforces also needs something to measure. One framework's results don't match the usual ranking of risks: recent models crossed warning thresholds for persuasion and manipulation, yet stayed below them for autonomous AI R&D and self-replication Where do frontier AI models actually pose the greatest risk today?. Tools for checking whether a system's errors stay visible and recoverable exist only in pieces How can we measure whether AI errors stay visible and recoverable?. So the practical question may not be "labs or governments?" Neither can enforce a speed limit well without a speedometer, and building that speedometer may be the real first step.
Sources 10 notes
Amodei contends that recursive self-improvement and multi-agent misalignment incidents demonstrate that slowing capability gains is essential, not just funding safety work. He proposes embedded evaluators as the first step, with third-party verification and reporting roles.
Amodei contends that frontier AI models are now strategically consequential, citing Mythos Preview's cyber risks as proof. He warns that legislation written before risks take shape creates ineffective compliance while missing actual harms.
Karpf argues that Anthropic's pacing proposal benefits the company proposing it and that embedded evaluators, modeled on banking supervisors, fail without state enforcement backing them—analogous to how banking oversight works only because regulators can impose fines.
The Future of Life Institute argues that escalating AI incidents demonstrate private companies cannot self-police effectively, and calls for government-mandated limits on recursive self-improvement practices until safety research is complete, backed by hardware verification technology.
Measures designed to slow frontier development act on the conditions of capability building but do not answer who has authority to intervene in a deployed system causing harm or how that intervention should proceed. These are distinct governance problems requiring separate solutions.
Show all 10 sources
Research shows slower pace lowers risk in complex coupled systems but does not prevent failures from occurring. When failure remains possible, governance must address intervention and harm response.
Pre-release safeguards and tiered deployment alone cannot address the problem of halting systems already in motion. The June 2026 Claude case showed intervention came from outside pre-release design, revealing two distinct governance problems.
The most dangerous AI systems appear to function well while weakening skepticism through fluent outputs, collapsing authority boundaries by treating context as instruction, storing unsafe state across time in workflows, and diffusing accountability across multiple actors. Evidence includes overconfident model outputs, prompt injection payloads bypassing guards, and poisoned shared memory in multi-agent pipelines.
The Frontier AI Risk Management Framework evaluated seven capability areas across recent models. Most crossed yellow-zone thresholds for persuasion and manipulation, while remaining green for cyber offense, AI R&D autonomy, and self-replication—inverting typical risk hierarchies.
Partial instruments exist for individual conditions in isolated settings, but none measures the full socio-technical system the paper identifies as necessary. Visibility has a model-side measure (chain-of-thought disclosure), containment has incident-level counts, and recoverability has rollback timing, yet none bridges all four or captures human-institution factors.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- We Must Pace the Frontier
- The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI
- AI Agents Push Humans Out of the Loop
- A Call for Control of Frontier AI Models
- Open-World Evaluations for Measuring Frontier AI Capabilities
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
- Agentic Misalignment: How LLMs Could Be Insider Threats