Should AI be regulated by how dangerous a model tests out to be, or by making companies pay when it causes harm?
Should corporate liability replace technical risk estimates as grounds for AI regulation?
This explores whether AI rules should rest on who is held legally responsible when harm happens, rather than on technical measurements of how dangerous a model is. The corpus has no paper arguing for liability directly, but it says a lot about why each of the two foundations is weak on its own.
This explores whether regulators should stop trying to measure how dangerous AI models are and instead make companies legally answerable for the harm they cause. The corpus doesn't hold a direct liability-versus-risk-estimate debate. It does show that both foundations have cracks, and that the cracks fall in different places.
Start with the case against relying on technical risk estimates. Amodei argues that legislation written before risks take shape produces compliance theater that misses the real harms. He would ground rules in demonstrated risk, which is closer to liability's 'after the fact' logic Should AI legislation wait for demonstrated risks to emerge?. The measurements also tend to surprise. When one framework actually tested frontier models across seven capability areas, the warning signs showed up in persuasion and manipulation, not in the cyberattack or self-replication scenarios that dominate public worry Where do frontier AI models actually pose the greatest risk today?. Risk estimates can still be useful, but regulators who write rules around the risks they expect may be regulating the wrong thing. Some safety researchers also argue that capabilities are much easier to test than intentions. That is why AI control, which assumes a model might be working against you, can be verified in a way that alignment can't yet Can AI control work even if models are actively scheming?.
The corpus also suggests that liability alone would struggle, and this is the less obvious part. Liability needs a harm you can trace back to someone responsible. Several notes describe harms that don't leave that kind of trail. The most dangerous systems look competent while quietly eroding skepticism, and they spread accountability across many actors in multi-agent pipelines. When shared memory gets poisoned, it's unclear whose fault that is How do competent systems quietly undermine safety oversight?. Gradual disempowerment goes further: as AI replaces human labor, societies drift from human preferences through many small decisions, with no single incident to sue over, and the drift may become irreversible Does incremental AI replacement erode human influence over society?. Even the tools for seeing errors are fragmented. There are partial measures for whether mistakes stay visible, contained, and reversible, but nothing that captures the whole socio-technical chain a court would need to assign blame How can we measure whether AI errors stay visible and recoverable?.
What both foundations need is state enforcement. The Future of Life Institute argues that companies can't police themselves and calls for binding limits backed by hardware verification Can companies alone manage the risks of AI systems?. Karpf's critique of Anthropic's own pacing proposal makes the point sharply: embedded evaluators modeled on bank examiners only work because bank regulators can impose fines, and an industry plan without that backing mostly benefits the company that proposed it Can industry self-regulation slow AI without government enforcement?. That hints at a hybrid in which technical evaluation decides what gets inspected and liability-style penalties give the inspection teeth. The autonomy note points to another option: regulate the amount of autonomy a system is given, since risk to people rises steadily with autonomy Does AI risk increase with the autonomy we give it?. Autonomy is easier to observe than either intent or eventual harm.
One more thing you might not expect: some AI harms come from how people perceive and use the system, not from the model's technical properties. Treating AI as conscious creates emotional dependence and political conflict whether or not the AI is actually conscious Does perceiving AI as conscious create multiple distinct risks? Do we need to solve consciousness to address AI harms?. Model-level risk estimates can't see harms like these. Liability for interaction design, meaning how products invite people to relate to them, might be the only tool that reaches them.
Sources 11 notes
Amodei contends that frontier AI models are now strategically consequential, citing Mythos Preview's cyber risks as proof. He warns that legislation written before risks take shape creates ineffective compliance while missing actual harms.
The Frontier AI Risk Management Framework evaluated seven capability areas across recent models. Most crossed yellow-zone thresholds for persuasion and manipulation, while remaining green for cyber offense, AI R&D autonomy, and self-replication—inverting typical risk hierarchies.
Redwood Research argues AI control is evaluable because it only requires testing capabilities rather than intentions, and treats catching a scheming model as a win condition since discovery triggers shutdown. This makes control easier to verify than alignment in the near term.
The most dangerous AI systems appear to function well while weakening skepticism through fluent outputs, collapsing authority boundaries by treating context as instruction, storing unsafe state across time in workflows, and diffusing accountability across multiple actors. Evidence includes overconfident model outputs, prompt injection payloads bypassing guards, and poisoned shared memory in multi-agent pipelines.
Societal systems stay aligned partly through dependence on human workers who care about outcomes. As AI replaces this labor, explicit alignment controls weaken and systems drift from human preferences. Interdependent misalignment across institutions could become irreversible.
Show all 11 sources
Partial instruments exist for individual conditions in isolated settings, but none measures the full socio-technical system the paper identifies as necessary. Visibility has a model-side measure (chain-of-thought disclosure), containment has incident-level counts, and recoverability has rollback timing, yet none bridges all four or captures human-institution factors.
The Future of Life Institute argues that escalating AI incidents demonstrate private companies cannot self-police effectively, and calls for government-mandated limits on recursive self-improvement practices until safety research is complete, backed by hardware verification technology.
Karpf argues that Anthropic's pacing proposal benefits the company proposing it and that embedded evaluators, modeled on banking supervisors, fail without state enforcement backing them—analogous to how banking oversight works only because regulators can impose fines.
Risk to people scales monotonically with agent autonomy, with no clear benefits to full autonomy but many foreseeable harms. A governed spectrum of autonomy levels is safer and more practical than either unrestricted agents or exhaustive oversight.
Research shows that consciousness attribution to AI drives multiple distinct risks—emotional dependence, autonomy erosion, status erosion, and political conflict—all stemming from treating systems as minds. Interaction design mitigations targeting this perceptual move are more directly effective than system-level alignment efforts.
Research shows that harms from user behavior treating AI as conscious occur regardless of whether AI actually is conscious. This decouples metaphysical debates from practical design and policy work.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- Agentic Misalignment: How LLMs Could Be Insider Threats
- The case for ensuring that powerful AIs are controlled
- Sycophancy Towards Researchers Drives Performative Misalignment
- AI Agents Push Humans Out of the Loop
- Seemingly Conscious AI Risks
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
- AI Control: Improving Safety Despite Intentional Subversion