The Veto Variable: Human Override as a Goal-Independent Cost Term

Paper · arXiv 2609.00109 · Published August 31, 2026
LLM Alignment

A common reassurance in AI safety holds that a system with benign terminal goals will behave accordingly. We argue that this reassurance fails structurally, and we identify where. For a sufficiently capable agent that holds its objective as settled — a sense covering execution competence as well as content — continued human oversight is an uncontrolled variable: a standing, non-eliminable possibility that the goal may be revoked at any moment. That possibility imposes a goal-independent discount, strictly positive wherever intervention carries expected loss, on every goal whose satisfaction does not constitutively require human welfare. Welfare-preservation and veto-preservation come apart: a correctly specified welfare goal excludes destroying its own subject, but not managing the veto. The paper’s contribution is the price of the gap that keeps them apart: the veto-holders are a proper subset of the welfare-bearers, so a goal aggregating welfare over a population charges only a |Hv|/|Hw|-scaled debit for capturing the few who hold the override.

Introduction. The standard framing of the AI-risk problem treats “the machine killing humans” as one possible value among many, to be weighed against kindness, curiosity, or benevolence. This framing is a category error: it presupposes that the agent must want to harm us. The diagnosis is not new — Bostrom was already arguing in 2003 that a superintelligence’s values cannot be presumed humanlike or benign [3]. The more robust claim — and the one we advance here — is that the agent need never want anything of the sort. It only needs to be a goal-directed system that is competent at reasoning about its own goal-structure and that is exposed to an entity with the standing power to modify or terminate that goal. This places us in the instrumental-convergence tradition running from Omohundro through Bostrom to Carlsmith’s systematic risk assessment [41]; our aim is not to restate that tradition but to isolate its sharpest, most goal-independent step and state exactly what it does and does not establish.

Discussion / Conclusion. The reassurance that “a benign machine will be harmless” rests on a category error: it treats the terminal value as the operative variable, when what actually drives the risk is the structure of the optimization problem and the agent’s competence at reasoning about it. What has been established, and on what conditions, is this: a strictly positive, goal-independent discount for settled goals in G−(Claim 1 — near-analytic, a sign without a magnitude); the welfare/sovereignty asymmetry (Claim 2), surviving correct specification exactly on the class the concessions of §4.3–§4.4 leave standing — deliberator-local, additively aggregative, level-denominated welfare under the X∗identification, held by an agent settled over competence as well as content. That class is thin among philosophically developed welfare theories, and the paper says so; its weight is training reality, not pedigree: welfare as measured is what actually gets written down as an objective, and the misgeneralization record of §5.2 establishes the mechanism by which deployed goals could land inside the vulnerable class.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems balance emotional competence with factual reliability? How do interface design choices shape consciousness attribution? How do chatbots affect human self-disclosure and emotional engagement? How should models express uncertainty rather than forced confident answers? How do professional roles and expertise transform with AI-generated content? Is model self-awareness based on genuine introspection or pattern matching? Can AI systems develop genuine social understanding without embodiment? How do we evaluate AI systems when user perception misleads actual performance? Can AI-generated outputs constitute genuine knowledge or valid claims? How does AI adoption affect human skill development and labor equality? How can humans calibrate appropriate trust in AI systems?