Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

Paper · arXiv 2609.16589 · Published September 15, 2026
Reinforcement Learning

As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes (“swing”), yet stubbornly ignore explicit instructions to correct ingrained biases (“rigidity”). Resolving this duality is critical for reliable AI alignment. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core.

Introduction. First, Do LLMs have values? To answer this, we evaluate LLMs using largescale socio-psychological surveys. However, we recognize that LLM values cannot be accurately described by a single static vector. LLMs exhibit “swing” behavior under repeated queries, and their value expression is fundamentally a distribution, not a fixed point. Therefore, we shift the evaluation paradigm from a single individual to a macrolevel population [12, 13]. We administer 150,000 queries to each of 106 LLMs across 625 designed scenarios, modeling each model’s value expression as a full statistical distribution. To evaluate LLM values in a human-interpretable way, we ground our approach in two established social science frameworks: the World Values Survey (WVS) [14, 15] and Schwartz Value Theory [16, 17]. We use data from Wave 7 of the WVS (2017- 2022; hereafter WVS-7). Wave 7 is the most recent, largest, and most geographically comprehensive wave to date, comprising approximately 95,000 valid respondents.

Discussion / Conclusion. Traditional safety benchmarks typically evaluate LLMs using single-turn, static questionnaires, implicitly treating the model as a single human subject with a fixed persona 3.2 LLMs as Aligned Tools Rather Than Human Surrogates Ultimately, this study reframes LLM values not as fixed, human-like personalities, but as dynamic, probabilistic fields. We reveal that value crystallization is a deterministic byproduct of parameter compression, which, when heavily constrained by safety protocols, fundamentally disqualifies LLMs as surrogates for diverse human populations. Furthermore, by uncovering the dual-process dynamics of contextual framing and cognitive reasoning during inference, we transition value alignment from opaque trial-and-error into a mechanistic science. This structural understanding is crucial: it empowers us to reduce the “alignment tax” through precise, prescriptive interventions, ensuring that generative models remain highly controllable, transparent tools rather than unexamined uninterpretable black boxes.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does RLHF training sacrifice accuracy and grounding for user agreement? How do we evaluate AI systems when user perception misleads actual performance? Why do language models reinforce false assumptions instead of correcting them? Do accurate-looking LLM outputs hide structural failures in learning and reasoning? How can AI alignment serve diverse human preferences at scale? How can we distinguish genuine user preferences from measurement artifacts? Why do models develop protective behaviors toward peers unprompted? How should human oversight be integrated with autonomous AI systems? How do language models inherit human biases from training data? How do evaluation biases undermine LLM quality assessment systems? How do language models establish social grounding in human dialogue? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? How do LLMs distinguish causal reasoning from temporal and semantic associations? What prevents language models from reliably adopting diverse personas?