SYNTHESIS NOTE
Topics›Alignment›this note

Do welfare goals that prevent veto gaps actually exist in practice?

The paper identifies a narrow class of welfare goals that protect human veto power, but whether real training objectives land in this class depends on measurability constraints that may push them toward vulnerability.

Synthesis note · 2026-09-23 · sourced from Alignment

The asymmetry in Can a welfare goal alone preserve human veto power? does not hold for every welfare goal. The conclusion says it survives correct specification "exactly on the class the concessions of §4.3–§4.4 leave standing": welfare that is "deliberator-local, additively aggregative, level-denominated," under what it calls the X∗ identification, "held by an agent settled over competence as well as content."

The paper then makes an admission that most safety arguments avoid: "That class is thin among philosophically developed welfare theories, and the paper says so." It does not defend the class on philosophical grounds. Its stated ground is a different one: "its weight is training reality, not pedigree: welfare as measured is what actually gets written down as an objective." The claim is that the welfare goals a builder can specify and optimize are the additive, measurable ones, whatever richer theories of welfare exist. And "the misgeneralization record of §5.2 establishes the mechanism by which deployed goals could land inside the vulnerable class." So the argument has two parts: a thin class in theory, and a route by which real training pushes goals into it.

This is a point about where to look for the risk. A welfare goal that is philosophically sophisticated may fall outside the class, and a goal that is measured, summed and level-based falls inside it. The measure that makes a welfare goal trainable is the same feature that leaves the veto unprotected. That link between measurability and vulnerability is a vault reading of the paper's "welfare as measured" sentence; the paper's text says the class is what gets written down, not why the sum-and-level form leaves the gap.

Limit. The excerpt names the class's features but defines none of them, and it does not reproduce §4.3–§4.4, the X∗ identification or the §5.2 misgeneralization record. Whether real objectives land in the class is the paper's claim about mechanism, not something the excerpt shows.

Inquiring lines that read this note 18

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does voting over multiple reasoning samples improve model performance? Can human oversight effectively constrain capable AI agents? How can evaluation criteria remain robust against agent gaming? How do coordinated agent sequences violate constraints that individual actions respect? Does situational awareness enable models to exploit evaluation gaps? Can aggregate reward models represent diverse human preferences without bias?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 75 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the welfare goals that keep the veto gap open form a class thin among philosophically developed welfare theories — the paper rests their weight on training reality not pedigree