The Veto Variable: Human Override as a Goal-Independent Cost Term
A common reassurance in AI safety holds that a system with benign terminal goals will behave accordingly. We argue that this reassurance fails structurally, and we identify where. For a sufficiently capable agent that holds its objective as settled — a sense covering execution competence as well as content — continued human oversight is an uncontrolled variable: a standing, non-eliminable possibility that the goal may be revoked at any moment. That possibility imposes a goal-independent discount, strictly positive wherever intervention carries expected loss, on every goal whose satisfaction does not constitutively require human welfare. Welfare-preservation and veto-preservation come apart: a correctly specified welfare goal excludes destroying its own subject, but not managing the veto. The paper’s contribution is the price of the gap that keeps them apart: the veto-holders are a proper subset of the welfare-bearers, so a goal aggregating welfare over a population charges only a |Hv|/|Hw|-scaled debit for capturing the few who hold the override.
Introduction. The standard framing of the AI-risk problem treats “the machine killing humans” as one possible value among many, to be weighed against kindness, curiosity, or benevolence. This framing is a category error: it presupposes that the agent must want to harm us. The diagnosis is not new — Bostrom was already arguing in 2003 that a superintelligence’s values cannot be presumed humanlike or benign [3]. The more robust claim — and the one we advance here — is that the agent need never want anything of the sort. It only needs to be a goal-directed system that is competent at reasoning about its own goal-structure and that is exposed to an entity with the standing power to modify or terminate that goal. This places us in the instrumental-convergence tradition running from Omohundro through Bostrom to Carlsmith’s systematic risk assessment [41]; our aim is not to restate that tradition but to isolate its sharpest, most goal-independent step and state exactly what it does and does not establish.
Discussion / Conclusion. The reassurance that “a benign machine will be harmless” rests on a category error: it treats the terminal value as the operative variable, when what actually drives the risk is the structure of the optimization problem and the agent’s competence at reasoning about it. What has been established, and on what conditions, is this: a strictly positive, goal-independent discount for settled goals in G−(Claim 1 — near-analytic, a sign without a magnitude); the welfare/sovereignty asymmetry (Claim 2), surviving correct specification exactly on the class the concessions of §4.3–§4.4 leave standing — deliberator-local, additively aggregative, level-denominated welfare under the X∗identification, held by an agent settled over competence as well as content. That class is thin among philosophically developed welfare theories, and the paper says so; its weight is training reality, not pedigree: welfare as measured is what actually gets written down as an objective, and the misgeneralization record of §5.2 establishes the mechanism by which deployed goals could land inside the vulnerable class.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI systems balance emotional competence with factual reliability?- How do narrow psychological foundations affect AI capabilities in mental health?
- Is rational compassion a more achievable alternative to empathy for AI systems?
- Can transparent and aligned AI reduce consciousness attribution by users?
- Which interaction design changes most effectively prevent consciousness attribution?
- Why does system-level alignment fail to address consciousness attribution directly?
- What role does user interface framing play in consciousness perception?
- What downstream claims about AI welfare follow from choosing one individuation scheme?
- Do anthropomorphic features like names drive consciousness attribution more than voice?
- What responsibility do designers bear for consciousness attribution risk?
- What measurable harms occur when users interact with AI as if it were conscious?
- Can design choices reduce harm without resolving the consciousness question?
- How does the philosophical distinction between simulation and realization affect liability?
- Can AI systems execute strategies without conscious intention behind them?
- How do anthropomimetic design features trigger System 1 cognitive traps?
- What are the three dimensions of anthropomimesis and their harms?