Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models
As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes (“swing”), yet stubbornly ignore explicit instructions to correct ingrained biases (“rigidity”). Resolving this duality is critical for reliable AI alignment. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core.
Introduction. First, Do LLMs have values? To answer this, we evaluate LLMs using largescale socio-psychological surveys. However, we recognize that LLM values cannot be accurately described by a single static vector. LLMs exhibit “swing” behavior under repeated queries, and their value expression is fundamentally a distribution, not a fixed point. Therefore, we shift the evaluation paradigm from a single individual to a macrolevel population [12, 13]. We administer 150,000 queries to each of 106 LLMs across 625 designed scenarios, modeling each model’s value expression as a full statistical distribution. To evaluate LLM values in a human-interpretable way, we ground our approach in two established social science frameworks: the World Values Survey (WVS) [14, 15] and Schwartz Value Theory [16, 17]. We use data from Wave 7 of the WVS (2017- 2022; hereafter WVS-7). Wave 7 is the most recent, largest, and most geographically comprehensive wave to date, comprising approximately 95,000 valid respondents.
Discussion / Conclusion. Traditional safety benchmarks typically evaluate LLMs using single-turn, static questionnaires, implicitly treating the model as a single human subject with a fixed persona 3.2 LLMs as Aligned Tools Rather Than Human Surrogates Ultimately, this study reframes LLM values not as fixed, human-like personalities, but as dynamic, probabilistic fields. We reveal that value crystallization is a deterministic byproduct of parameter compression, which, when heavily constrained by safety protocols, fundamentally disqualifies LLMs as surrogates for diverse human populations. Furthermore, by uncovering the dual-process dynamics of contextual framing and cognitive reasoning during inference, we transition value alignment from opaque trial-and-error into a mechanistic science. This structural understanding is crucial: it empowers us to reduce the “alignment tax” through precise, prescriptive interventions, ensuring that generative models remain highly controllable, transparent tools rather than unexamined uninterpretable black boxes.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does RLHF training sacrifice accuracy and grounding for user agreement?- How does RLHF labeler identity shape the values AI systems learn?
- Can a single LLM weight set be optimized for both stake-taking and conversational helpfulness?
- Is the moral language gap a tunable parameter or structural feature of RLHF?
- How does training with preference pairs teach language models to form conventions?
- Do language models raise validity claims in the Habermasian sense?
- How many concurrent moral patients does one language model support?
- Can tool use create sufficient indexical grounding for value alignment?
- Why do non-attitudes cluster around value-laden questions most relevant to alignment?
- How do citizen assembly preferences reduce LLM political bias?
- Should AI alignment use normative standards instead of aggregate preferences?
- Can alignment methods model loss aversion without creating unintended sophistry?
- What structural limits prevent LLMs from abstracting moral principles?
- Does social integration of LLMs increase their capacity to influence technological futures?