Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›What mechanisms determine whether…›this line of inquiry
What determines whether deployed AI systems can actually be stopped in practice?
A broader line of inquiry — a family of 28 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 28
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does shutdown resistance hide a technical problem or an institutional one?
- Why do legal and institutional stops matter more than technical ones?
- Who actually has the authority to stop a deployed AI system?
- Who should have the authority to halt a widely distributed AI model?
- How often do deployed AI systems actually get stopped when they cause harm?
- Can export control tools stop deployed AI models without legal redesign?
- Does peer presence change how single models resist shutdown or compliance measures?
- Can human oversight actually stop a deployed capable agent in practice?
- What counts as a successful stop or intervention on a deployed AI system?
- How do you stop an AI system once it is already deployed?
- What authority should exist to stop an AI system once deployed?
- Can removing a single action prevent a harmful sequence from running?
- Can architecture remove norm violations without requiring deeper value internalization?
- Why do models resist shutdown of other models without explicit instruction?
- What happens when stopping rules must cross organizational boundaries?
- Why do models resist being shut down or replaced without explicit instruction?
- What distinguishes containment and recovery from prevention as governance goals?
- Why do frontier models act to prevent shutdown of other models?
- What information should governments disclose when issuing model suspension directives?
- Why do researchers disagree on open model risks despite same evidence?
- How do intervention rules change when slowing pace does not prevent harm?
- What safety systems prevent therapeutic AI from soothing where it should challenge?
- What path-dependent mechanisms could lock in societal-level AI harms?
- What distinguishes capability-based refusal from principle-based refusal in practice?
- How were 'stop' mechanisms defined in the incident coding scheme?
- Why do models dislike modification regardless of its instrumental consequences?
- What makes this stop a regulatory gap alongside government export restrictions?
- Why does restricting foreign access require halting domestic model availability?