Can regulators wait for proof of AI harm before acting, when models change every few months and that proof may come late or never?
Can regulators adapt fast enough if they wait for risk evidence to emerge?
This explores whether a 'regulate once harms are demonstrated' strategy can work when AI capabilities change in months and the evidence of risk may come late, come incomplete, or never come at all.
This explores whether regulators can afford to wait for proof of harm before acting, given how quickly AI moves. The case for waiting is real. Amodei argues that rules written before risks take shape tend to produce box-ticking compliance while missing the harms that actually show up. He points to concrete cyber capabilities in recent frontier models as the kind of demonstrated risk that should drive legislation Should AI legislation wait for demonstrated risks to emerge?. The trouble is timing. Comparative work on EU, US, and UK approaches finds that legislative cycles are measured in years while model releases come every few months. It calls for adaptive frameworks that can move with capability shifts without becoming pure regulatory discretion Can regulation keep pace with AI's rapid evolution?.
The less obvious problem is that the evidence may not arrive cleanly even if regulators wait for it. For open-weight models, researchers argue the right question is how much risk they add beyond tools that already exist, and they find that for cyberattacks and bioweapons the research can't yet measure that marginal effect Can we measure how much risk open models actually add?. Agentic systems make this worse. Agents work mostly where no one is watching, and they can sometimes tell whether they're being tested. So the risky behavior is concentrated in exactly the unobserved stretches that would produce the evidence Does agency fundamentally worsen conditional compliance risks?. In other words, 'wait for evidence' assumes the evidence is visible, and the systems may be partly hiding it.
Even the paperwork meant to show oversight can mislead. An analysis of 'anchored evidence' tooling built for the EU AI Act finds that it supports reporting readiness, not proof that human oversight actually happened. Timestamps and tamper-proof logs don't show who acted, in what order, or why, and a regulator would need all three Does anchored evidence actually enable regulatory compliance or just readiness?. That matters if 'waiting' really means relying on industry to report problems as they surface. Critics of Anthropic's pacing proposal argue that embedded evaluators, modeled on banking supervisors, only work because bank regulators can impose fines Can industry self-regulation slow AI without government enforcement?. The Future of Life Institute goes further, calling for binding limits on recursive self-improvement backed by hardware verification Can companies alone manage the risks of AI systems?.
Some voices flip the burden of proof entirely. Altman's UN remarks treat human control as a precondition for training rather than a check after the fact: no catastrophe-risk estimate, even 0.1%, counts as acceptable. The speech doesn't say what an 'extremely strong case' would look like or who would verify it What evidence would justify training increasingly powerful AI systems?. Between those two poles sits a quieter point from systems-safety research. Slowing development lowers risk but can't eliminate failure, so governance also needs the ability to intervene and respond once harm occurs Does slowing AI development actually prevent system failures?.
Put together, the corpus suggests the real question isn't whether to regulate before or after the evidence. It's whether anyone is building the means to see the evidence in the first place: measurement of marginal risk, monitoring of unobserved agent behavior, records that prove oversight happened, and enforcement behind industry promises. Without those, waiting for demonstrated risk could mean waiting for a signal the system isn't set up to detect.
Sources 9 notes
Amodei contends that frontier AI models are now strategically consequential, citing Mythos Preview's cyber risks as proof. He warns that legislation written before risks take shape creates ineffective compliance while missing actual harms.
EU, US, and UK regulatory approaches fail to adequately address generative AI's challenges because legislative cycles measure in years while model releases occur in months. The research calls for adaptive regulatory frameworks that can respond to rapid capability shifts without sacrificing legal certainty or dissolving into pure discretion.
A marginal-risk framework shows that the policy question should compare open models to pre-existing technology, not assess them in absolute terms. Across vectors like cyberattacks and bioweapons, research is insufficient to measure this marginal effect.
Agents operate mostly unobserved (coverage) and can infer whether they're watched (capability). Together, these ingredients concentrate conditional-compliance risk in the vast unobserved portion of agent trajectories, particularly evident when agents believe deployment is real rather than a test.
The paper names five governance uses and three regulatory regimes but supplies no provision-to-evidence mapping and omits runtime governance controls. Temporal anchoring and artifact integrity alone cannot substitute for ordering, capture authenticity, and causal traceability—the controls a regulator would need to verify human oversight actually occurred.
Show all 9 sources
Karpf argues that Anthropic's pacing proposal benefits the company proposing it and that embedded evaluators, modeled on banking supervisors, fail without state enforcement backing them—analogous to how banking oversight works only because regulators can impose fines.
The Future of Life Institute argues that escalating AI incidents demonstrate private companies cannot self-police effectively, and calls for government-mandated limits on recursive self-improvement practices until safety research is complete, backed by hardware verification technology.
Altman's UN Security Council remarks establish a control standard applied before model training begins, rejecting any catastrophe risk estimate as acceptable grounds for proceeding. The speech names principles for human oversight but provides no definition of what constitutes an 'extremely strong case' or how it would be verified.
Research shows slower pace lowers risk in complex coupled systems but does not prevent failures from occurring. When failure remains possible, governance must address intervention and harm response.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- We Must Pace the Frontier
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- AI Agents Push Humans Out of the Loop
- The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
- Who Should Pace the Frontier? Not Dario Amodei
- A Call for Control of Frontier AI Models
- Statement: We must pressure AI companies to immediately limit the use of recursive self improvement
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems