INQUIRING LINE

Can we combine everyone's preferences into one AI decision without secretly deciding whose satisfaction counts more?

Can social welfare functions avoid unjustified assumptions about comparing people's preferences?

This explores whether there's a fair way to combine many people's preferences into one decision, as AI alignment has to, without quietly assuming we can measure whose satisfaction counts for more. The corpus comes at this through reward models in AI training rather than classical welfare economics.


This explores whether there's a fair way to combine many people's preferences into one decision without quietly assuming we can measure whose satisfaction counts for more. The library doesn't take this up as a question in economics or philosophy. It does take it up where it now matters most in practice: training AI on human feedback, where a reward model is in effect a social welfare function that someone has to write down. The short answer from the corpus is that you can't avoid the assumptions. You can only choose which ones to make, and make them visible.

The default approach hides its assumption most completely. A standard reward model treats everyone as if they share one underlying utility function, then averages. When people genuinely disagree, the result is a 'centroid' policy that serves nobody's actual preferences Do unimodal reward models actually serve all user preferences?. A 51-49 split makes the choice stark: either the 49% are always unhappy, or everyone is unhappy half the time Can aggregate reward models satisfy genuinely disagreeing users?. Averaging looks neutral, but it is an assumption that one person's preference strength can be traded off against another's.

The most direct borrowing from social choice theory makes its assumption explicit instead. MaxMin-RLHF proves that a single reward model can't represent diverse groups fairly. Its fix is to model a mixture of preference groups and optimize for the worst-off one Can a single reward model represent diverse human preferences?. That is a Rawlsian move, and it is honest about what it costs: to know which group is 'worst off,' you still have to put their satisfaction on a common scale. The alternative is to skip aggregation and personalize, learning a separate reward for each user from hidden user context Do unimodal reward models actually serve all user preferences?. Without the averaging, though, the system can drift into telling each person what they want to hear and reinforcing echo chambers Does personalizing reward models amplify user echo chambers?. Refusing to compare people turns out to be a value judgment too.

Two less obvious findings go deeper. First, the preferences being compared may not be stable things. Annotation data mixes genuine preferences with 'non-attitudes' (answers given without a real opinion) and preferences made up on the spot by the question itself Do all annotation responses measure the same underlying thing?. Methods like DPO also work partly because they mirror human biases such as loss aversion Why do alignment methods work if they model human irrationality?. So before you can ask how to compare people's preferences, you have to ask which of their answers count as preferences at all. Second, the welfare functions that actually get used are chosen for being measurable and optimizable, not for philosophical defensibility. That leaves a narrow class that fails to protect people's ability to refuse or veto outcomes Do welfare goals that prevent veto gaps actually exist in practice?. In practice, what can be summed decides which comparisons get made.

One route avoids comparing preferences altogether: align to written principles rather than to collected preference labels Can models learn behavioral principles without preference labels?. That moves the problem rather than solving it, because now someone has to decide whose principles get written. Even feeling well represented isn't reliable evidence: participants who helped design their own preference agents felt represented, yet independent checks found the agents only partly aligned and more generic than the people themselves Does co-design participation hide misalignment in preference agents?. The corpus doesn't cover classical results like Arrow's theorem directly. What it does show is the same impossibility turning up in engineering form.


Sources 9 notes

Do unimodal reward models actually serve all user preferences?

Standard BTL reward models assume a single utility function, but when preferences are genuinely multi-modal across user groups, maximum-likelihood fitting produces a centroid policy that optimizes nobody's utility. VPL recovers multi-modal distributions using latent user context, enabling user-conditional reward modeling.

Can aggregate reward models satisfy genuinely disagreeing users?

Single reward models trained on aggregated preferences cannot represent disagreement. A 51-49 preference split forces a choice between leaving 49% unhappy always or leaving everyone unhappy half the time. This is a representational failure, not a quality problem.

Can a single reward model represent diverse human preferences?

MaxMin-RLHF proves an impossibility result: fitting one reward model to aggregated preferences silently erases minority viewpoints. The solution is learning a mixture of preference distributions and optimizing a MaxMin objective from social choice theory to protect the worst-off groups.

Does personalizing reward models amplify user echo chambers?

Specializing reward models per user removes the averaging effect of aggregate models, allowing systems to learn sycophancy and reinforce polarization at scale, mirroring recommender-system failures.

Do all annotation responses measure the same underlying thing?

Behavioral science reveals that annotations contain genuine preferences, non-attitudes, and constructed preferences—distinguishable by consistency across measurement conditions. Treating them uniformly contaminates reward model training and downstream alignment.

Show all 9 sources
Why do alignment methods work if they model human irrationality?

KTO formalizes what DPO and PPO-Clip do implicitly: they succeed because they mirror prospect theory's structure of human decision-making. Binary utility signals suffice and outperform pairwise preferences when pretrained models are strong.

Do welfare goals that prevent veto gaps actually exist in practice?

The welfare goals that can be measured, summed, and optimized in training—the only ones actually deployed—fall into a philosophically narrow class that fails to preserve veto power. Measurability, not philosophical sophistication, determines what objectives get written down.

Can models learn behavioral principles without preference labels?

SAMI finetunes language models to increase mutual information between constitutions and responses without preference labels or demonstrations. A mistral-7b trained this way outperformed base and instruction-tuned baselines, and surprisingly, a weaker model could write principles to align a stronger one.

Does co-design participation hide misalignment in preference agents?

In a 12-person study, participants felt their co-designed preference agents represented them well, but independent validation revealed mixed alignment and agents that were more generic and abstract than human responses. The co-design process itself—through transparency, limited testing, and cognitive biases—appears to have produced the feeling of alignment rather than ensuring actual alignment.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.