Line of inquiry
Inquiring lines›What enables authentic and grounde…›What architectural and training st…›this line of inquiry
How can AI alignment serve diverse human preferences at scale?
A broader line of inquiry — a family of 35 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 35
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Should AI alignment use normative standards instead of aggregate preferences?
- Does a single LLM judge capture diverse human preferences in alignment training?
- Can communication problems and optimization problems be addressed with the same alignment approaches?
- Can AI-assisted alignment eventually solve fairness at scale?
- Should LLMs align with social roles instead of individual preferences?
- Why do standard preference alignment methods fail at the individual user level?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?
- Can alignment procedures be redesigned to serve multiple preference groups?
- Can bidirectional model updating between humans and AI reduce misalignment?
- Can alignment methods like DPO exploit or correct these surface feature biases?
- Why do text-based user summaries outperform embedding vectors for pluralistic alignment?
- Why do alignment values become problematic as language models scale?
- Can preference optimization and faithfulness measurement coexist as separate alignment objectives?
- Why does AI alignment fail when goals lack indexical grounding in values?
- Can tool use create sufficient indexical grounding for value alignment?
- Can a single AI system optimize multiple alignment dimensions simultaneously?
- What makes principle-response mutual information sufficient for behavioral alignment?
- Can preference trees structure alignment data for domains beyond math and code?
- Can constitutional AI alignment work without preference labels by maximizing input-response mutual information?
- How much does forcing single-choice answers damage alignment with complex intent?
- What preference optimization strategy works best for multi-turn social alignment?
- How does constitutional alignment compare to RLHF in removing human annotation costs?
- How should product specifications measure alignment without naming the dimension?
- Can alignment methods model loss aversion without creating unintended sophistry?
- How can developers balance multiple conflicting fairness goals simultaneously?
- How should historical preferences be weighted when users change their stated intent?
- What does egalitarian social choice theory contribute to AI alignment?
- What quality of curated data is minimally sufficient for alignment?
- What preference data do different personalized alignment methods actually need?
- Why do non-attitudes cluster around value-laden questions most relevant to alignment?
- Does DPO improve or harm LLM behavior in different training contexts?
- Which application domains like healthcare and education lack alignment research?
- What prevents human-centered objectives from being applied universally across all contexts?
- How do citizen assembly preferences reduce LLM political bias?
- Why does fixing harm require stakeholder input rather than universal developer definitions?