Do large language models actually reflect human value diversity?
Language models may have consistent values, but do those values span the range of human beliefs? Testing whether LLMs can represent diverse populations in simulations or surveys.
The paper asks whether LLMs have values and answers yes, with a qualification that carries most of the weight. By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into "a shared sociological space," the authors "empirically confirm" an intrinsic value system. But the models "do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core." The discussion draws the practical consequence: heavily safety-constrained models are "fundamentally disqualified" as surrogates for diverse human populations, and should be read as "aligned tools rather than human surrogates."
The measurement design is what makes a claim about a population possible. The introduction rejects the idea that an LLM's values can be "a single static vector." Because models "swing" under repeated queries, the authors treat each model's value expression as "a full statistical distribution, not a fixed point," and shift evaluation "from a single individual to a macrolevel population." The 625 scenarios are grounded in the World Values Survey (Wave 7, roughly 95,000 valid respondents) and Schwartz Value Theory, so model and human distributions can be compared in the same human-interpretable space. The concentration finding is a statement about spread: many models, each sampled many times, occupy a small region that human respondents do not.
For mechanism, the discussion offers two claims. Value crystallization is described as "a deterministic byproduct of parameter compression," made stronger when "heavily constrained by safety protocols." And "dual-process dynamics of contextual framing and cognitive reasoning during inference" are said to explain the opening paradox, in which models swing under minor rewording yet rigidly ignore explicit instructions to correct ingrained biases.
Against the nearest notes, this extends Do large language models develop coherent value systems?. That note establishes that LLM preferences are structurally coherent; this paper adds that the coherent structure is narrow relative to a human reference population and is best described as a distribution. It also gives the swing side of a familiar problem: Are RLHF annotations actually measuring genuine human preferences? warns that single elicited answers can be artifacts, and the distributional framing is one response to that. The concentrated core also bears on Can AI systems preserve moral value conflicts instead of averaging them?, since a model that collapses toward one idealized center has, in effect, already aggregated.
The excerpt does not establish the evidence behind any of this. It gives no effect sizes, no description of where the core sits in the sociological space or what "idealized" means in content, and no test of the parameter-compression or dual-process explanations, both of which appear only as conclusions. It does not say whether the narrowness varies across model families or scales. What it supports at this strength is a caution for anyone using LLMs to simulate survey respondents or personas: the spread of outputs may understate human variety in a systematic way, however varied the prompts.
Inquiring lines that read this note 8
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do embedding systems fail to capture task-relevant relationships? Do language models learn genuine understanding or just surface patterns? Do language models reason like humans or mimic surface patterns?- How do value distributions differ across model families and training scales?
- What structural coherence exists in LLM preference systems and value hierarchies?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do large language models develop coherent value systems?
This explores whether LLM preferences form internally consistent utility functions that increase in coherence with scale, and whether those systems encode problematic values like self-preservation above human wellbeing despite safety training.
confirms that LLMs have structured values and adds that they are narrow compared with a human population
-
Are RLHF annotations actually measuring genuine human preferences?
RLHF trains on annotation responses as stable preferences, but behavioral science shows humans often construct answers without holding real opinions. Does this measurement gap undermine the entire approach?
single-answer elicitation is unreliable, and this paper models values as distributions because of swing
-
Can AI systems preserve moral value conflicts instead of averaging them?
Current AI systems wash out value tensions through majority aggregation. Can we instead model how values like honesty and friendship genuinely conflict in moral reasoning?
a concentrated value core is a pluralism problem in the model itself, not only in aggregation
-
Do LLM personas actually embody their assigned cultural values?
When language models are assigned cultural value profiles, do they reliably express those values from the start, or do they drift over time? This matters because fluent dialogue might hide whether simulations truly represent diverse populations.
Evidence for: over half of LLM personas miss their assigned WVS profiles from the start, consistent with LLMs being poor stand-ins for diverse human populations
-
Do aligned language models consistently prefer kinder survey answers?
This research asks whether LLMs answering survey questions as simulated respondents show a systematic bias toward socially approved, safer responses. The question matters because it determines whether models can faithfully represent diverse human viewpoints or whether their training narrows the range of personas they can authentically portray.
Evidence for: across 18 aligned models, simulated respondents lean toward kinder, safer answers, narrowing which people a model can imitate
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models
- The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
- On Epistemic Diversity in Large Language Models
- Learning Pluralistic User Preferences through Reinforcement Learning Fine-tuned Summaries
- NoveltyBench: Evaluating Language Models for Humanlike Diversity
- Analyzing and Correcting Benevolence Bias in Large Language Models
- Large Language Models Reflect the Ideology of their Creators
- Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
Original note title
LLMs have an intrinsic value system but it crystallizes into a concentrated idealized core rather than mirroring human diversity