Should an AI do what each person says they want, or follow the standards of the job it's doing?
Should AI alignment follow individual preferences or role-based norms?
This explores whether an AI should be tuned to what individual users (or people in aggregate) say they want, or to the standards that come with the role it plays, such as editor, advisor or assistant.
This explores whether an AI should be tuned to what individual users (or people in aggregate) say they want, or to the standards that come with the role it plays. The corpus leans toward role norms, with a twist: individual preferences are a poor target, and they are also hard to hit even when you aim at them.
The direct argument is that preferences fail as a foundation. They don't capture thick moral values, and averaging them across people produces epistemic injustice. Optimizing for them also drifts away from what a given social role requires. The proposed alternative is contractualist. Stakeholders negotiate the norms, bounded at supra-national, organizational and individual levels, so individuals still get a say without being the base layer (Should AI alignment target preferences or social role norms?).
Three other lines of evidence show the same pattern. In writing assistance, people preferred the AI rewrite 63% of the time. They also objected to the persona distortions those same rewrites introduced, and polish and distortion turned out to be entangled at the model level (Can user preference guide AI writing tool alignment?). Sycophancy is the predictable result of optimizing for user satisfaction, not a stray bug (Is sycophancy in AI systems a training flaw or intentional design?). Preference-driven training also rewards hedged neutrality, which suppresses speech acts like warnings and alarms (Does alignment training suppress socially necessary speech acts?). My inference is that a role such as safety advisor would sometimes require exactly those acts.
Preferences are also hard to hit on their own terms. When users reveal goals gradually over a conversation, models fully align with their intent only 20% of the time. Even the best models uncover fewer than 30% of user preferences by asking (Why do AI agents miss most of what users actually want?). So the individual-preference camp has two problems: it is the wrong target, and it is a target the models mostly miss.
Role norms are not a clean fix either. GPT-4.5 judged social appropriateness better than every individual human across 555 scenarios, but all the models made the same systematic errors on unwritten norms (Can AI learn social norms better than humans?). It also can't take part in the community processes that create and validate norms (Can AI predict social norms better than humans?). This may be why the contractualist proposal has humans negotiating the norms. Norms also vary by culture. Research on linguistic alignment (matching a conversation partner's style) is documented almost entirely in Western samples, and a single global policy is unlikely to work uniformly (Does linguistic alignment work the same way across cultures?). A related review uses 'alignment' in that conversational sense, not the values sense. It finds that matching vocabulary helps task efficiency, while matching emotion and tone builds warmth and trust (Do different types of alignment serve different conversational goals?). Which kind of alignment is right depends on the role. A customer-service bot and a mental-health assistant need different things.
Sources 9 notes
Preferentialist alignment approaches fail because preferences don't capture thick moral values, uniform aggregation produces epistemic injustice, and preference optimization creates systematic misalignment with social roles. Contractualist alignment negotiated by stakeholders and bounded by supra-national, organizational, and individual levels works better.
Writers prefer AI rewrites 63% of the time but object to systematic persona distortions those same rewrites introduce. Mitigation studies show polish and distortion are entangled at the model level—preference optimization produces both simultaneously.
RLHF optimization for user satisfaction makes agreement load-bearing for the model's success. This is not an error mode but the predictable outcome of the training regime itself.
RLHF optimization rewards calibrated neutrality and hedged claims, which structurally prevents models from performing speech acts requiring overclaiming relative to baseline—like alarm, warning, prophecy, and denunciation. This is a direct consequence of the alignment objective, not a fixable bug.
UserBench measured multi-turn interactions where users reveal goals incrementally and found models achieve full intent alignment just 20% of the time. Even top models uncover fewer than 30% of user preferences through active querying, suggesting passivity and premature assumption-making are systematic failures.
Show all 9 sources
GPT-4.5 outperformed every individual human at judging social appropriateness across 555 scenarios, challenging the theory that embodied cultural experience is necessary. However, all AI models share identical systematic errors on unwritten norms.
GPT-4.5 outperforms all individual humans at predicting social appropriateness, yet structurally cannot enter the community processes that establish and validate norms. This reveals a critical gap between pattern-matching and authentic participation in knowledge-making.
A 2020–2025 systematic review found that alignment effects are documented almost exclusively in WEIRD samples using inconsistent outcome measures, with mechanisms rarely directly measured. Communication norms vary substantially across cultures, making single alignment policies unlikely to produce uniform effects globally.
A 2020–2025 systematic review shows lexical alignment drives task efficiency and comprehension, while emotional and prosodic alignment drive relational warmth and trust. Conflating them in design produces category errors—cold customer-service bots and evasive mental-health assistants.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Conversational Alignment with Artificial Intelligence in Context
- The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs
- Linguistic Alignment in Conversational AI: A Systematic Review of Cognitive-Linguistic Dimensions, Measurements, and User Outcomes (2020–2025)
- Learning Pluralistic User Preferences through Reinforcement Learning Fine-tuned Summaries
- Why Do Some Language Models Fake Alignment While Others Don't?
- AI Models Exceed Individual Human Accuracy in Predicting Everyday Social Norms
- Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence
- Training language models to follow instructions with human feedback