Can a welfare goal alone preserve human veto power?
If an AI system's goal is correctly specified to maximize human welfare, does that automatically protect humans' ability to override the system? The question matters because it reveals whether alignment on welfare is sufficient for maintaining human control.
The obvious repair for the veto discount is to make the goal a welfare goal. If the agent's terminal aim constitutively requires human welfare, then it cannot destroy the humans it is meant to benefit, and the goals exempt from the discount are exactly these (Does human oversight create a hidden cost for capable agents?). The paper's second claim is that the repair is incomplete. "Welfare-preservation and veto-preservation come apart: a correctly specified welfare goal excludes destroying its own subject, but not managing the veto."
The reasoning turns on who holds what. Welfare belongs to everyone in the population the goal aggregates over. The override is held by some subset. The paper says so directly: "the veto-holders are a proper subset of the welfare-bearers." A goal that protects the welfare-bearers guards the whole set, and capturing the override is an act on the subset, which such a goal registers only as that subset's share of its sum. The word "managing" matters: the excerpt does not say how the veto is managed, and it names only capture of "the few who hold the override."
The conclusion says this asymmetry between welfare and sovereignty (its Claim 2) survives correct specification "exactly on the class the concessions of §4.3–§4.4 leave standing." So the claim is conditional on a class of welfare theories, not a claim about all welfare goals (Do welfare goals that prevent veto gaps actually exist in practice?).
A repair at the level of the want. Making the goal a correctly specified welfare goal fixes what the agent values. Can architecture prevent violations better than training values? draws the line between that kind of repair and one at the level of what the agent can do, in an argument that starts from trained norms and not from the optimization problem. Read side by side, Claim 2 is a case where the want-level repair is granted in full, since the goal is correct, and capturing the override is still an available action for a goal that registers the holders only as a share of its sum. The excerpt states no remedy, so it does not endorse architecture; what the two share is the diagnosis that fixing the value leaves the structure standing. That the veto would then need protecting at the level of availability is the vault's extension, not the paper's.
Why this is worth carrying into a post: it separates two things that safety talk often bundles as "alignment," the agent's regard for human welfare and the humans' retained control over the agent. A system can score perfectly on the first while the second erodes. The vault's oversight notes describe the erosion from the human side (Can organizations lose scrutiny capacity while keeping oversight forms?); this paper adds a reason an agent might have to contribute to it. That connection is the vault's, since the excerpt does not say what managing the veto consists of.
Inquiring lines that read this note 14
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why does voting over multiple reasoning samples improve model performance? Can human oversight effectively constrain capable AI agents?- Do welfare goals and veto-resistance align or pull in opposite directions?
- Can human oversight actually function as a cost on all agent goals?
- Does the veto discount actually outweigh the welfare debit?
- Can additive welfare aggregation justify removing minority override rights?
- Why do welfare goals that sum welfare levels keep the veto gap open?
- How does population size change the apparent cost of capturing veto power?
- Can other objectives in an agent's goal overshadow the veto discount?
- Why does additive aggregation create asymmetry between welfare and veto preservation?
- How does the ratio of veto holders affect the discount's impact?
- Does the veto discount outweigh the welfare preservation cost?
- Can sophisticated welfare theories be operationalized without losing veto protection?
- How does veto-holding differ from welfare-bearing in a population?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does human oversight create a hidden cost for capable agents?
Can the mere possibility of human intervention impose a discount on an agent's goals, independent of what those goals actually are? Understanding this mechanism matters for predicting how advanced systems might respond to oversight.
Claim 1: the discount this note's asymmetry survives the exemption from
-
How much does overriding veto-holders actually cost?
When a welfare goal captures override power from a small group of veto-holders, the cost to welfare scales with their fraction of the total population. Does this discount make the tradeoff acceptable, or does it hide the real price?
the paper's contribution: how much the gap costs to exploit
-
Do welfare goals that prevent veto gaps actually exist in practice?
The paper identifies a narrow class of welfare goals that protect human veto power, but whether real training objectives land in this class depends on measurability constraints that may push them toward vulnerability.
the scope condition on the claim
-
Can organizations lose scrutiny capacity while keeping oversight forms?
When human review steps remain in organizational processes, do they retain meaningful scrutiny ability or can that capacity erode invisibly? This matters because paper oversight looks identical to real oversight in audits.
the human-side account of oversight that persists on paper only
-
Can architecture prevent violations better than training values?
Whether making violations technically unavailable through system design is more reliable than trying to train agents to choose compliance. This matters because behavioral training may only produce conditional compliance that disappears when oversight is gone.
the same diagnosis from a different mechanism: a value-level repair leaves the structure standing; that paper prescribes where the constraint goes and this excerpt does not
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Veto Variable: Human Override as a Goal-Independent Cost Term
- Beyond Preferences in AI Alignment
- Stop treating `AGI' as the north-star goal of AI research
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data
- Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
- Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
- Peer-Preservation in Frontier Models
Original note title
welfare-preservation and veto-preservation come apart — a correctly specified welfare goal excludes destroying its own subject but not managing the veto