Do LLM personas actually embody their assigned cultural values?
When language models are assigned cultural value profiles, do they reliably express those values from the start, or do they drift over time? This matters because fluent dialogue might hide whether simulations truly represent diverse populations.
The paper builds a World Values Survey (WVS)-grounded simulation in which culturally diverse agents with different communication styles hold longitudinal, value-laden discussions. Across approximately 4,000 conversations, 1,200 personas, 15 topics, and three models (GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B), it reports that "more than 50% of personas fail to express their assigned WVS profiles from the outset, while 2-7% drift after repeated conversations." The title states the ordering plainly: the failure happens before the drift. The discussion adds the practical warning that "conversational plausibility can mask failures of representational fidelity."
The paper frames validity as two requirements, that agents can "reliably instantiate and sustain" culturally grounded value profiles during interaction. Most of the measured failure sits in the first requirement, instantiation, and only a small share in the second, sustaining. The ablation supports this reading. Removing demographic details improves faithfulness for some models but "does not change the broader trend," since simulated value distributions "still systematically deviate" from the assigned profiles. The realism comparison points the same way: against human discussions, simulated dialogues trade stylistic consistency against semantic diversity differently, "often producing content-wise varied but stylistically repetitive exchanges." Fluent, varied surface text is what makes the value failure hard to see. From this the authors conclude that richer demographic prompting is not a sufficient safeguard and that such simulations should be treated as model-mediated instruments requiring validation against the populations they claim to represent.
Against the nearest notes, this extends How do we generate realistic personas at population scale? from demographic joint distributions to values: even when a persona is specified, the value profile may not come through. It also sharpens Why do LLM persona prompts produce inconsistent outputs across runs?. That note describes noise swamping persona signal, while this paper describes value distributions that deviate systematically. It qualifies optimism such as Can AI agents learn people better from interviews than surveys? and Can AI personas reliably replicate human experiment results?: success on individual responses or robust main effects does not show that an agent holds a stable value system inside a conversation.
The excerpt is silent on several points a reader would need. It does not say how faithfulness was scored, how the over-50% figure divides across the three models, which models improved when demographics were removed, or what a human baseline for drift would be. It offers no mechanism for why instantiation fails. The WEIRD-overgeneralization and algorithmic-monoculture points are stated as risks ("may therefore", "might get amplified"), not measured results, and the effect on marginalized groups is likewise a stated concern. What follows at this strength is narrow: a longitudinal drift check cannot stand in for a check at the start of the run, and a persona's values need validating against the target population before its conversations are used as evidence about that population.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do language models resist personality conditioning through prompts?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How do we generate realistic personas at population scale?
Current LLM-based persona generation relies on ad hoc methods that fail to capture real-world population distributions. The challenge is reconstructing the joint correlations between demographic, psychographic, and behavioral attributes from fragmented data.
extends population-level calibration failures from demographic distributions to the value profiles assigned to individual personas
-
Why do LLM persona prompts produce inconsistent outputs across runs?
Can language models reliably simulate different social perspectives through persona prompting, or does their run-to-run variance indicate they lack stable group-specific knowledge? This matters for whether LLMs can approximate human disagreement in annotation tasks.
contrast: that note finds run-to-run noise swamping persona signal, this one finds systematic deviation from assigned profiles
-
Can AI agents learn people better from interviews than surveys?
Can rich interview transcripts seed more accurate generative agents than demographic data or survey responses? This matters because it challenges how we build digital simulations of real people.
qualifies: individual-response fidelity does not show stable value expression in multi-turn conversation
-
Can AI personas reliably replicate human experiment results?
Exploring whether LLM-based persona simulations accurately reproduce experimental findings from published psychology and marketing research, and what factors determine when they succeed or fail.
qualifies: robust main effects can replicate while assigned value profiles still fail to appear
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
- Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models
- Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
- When Persona Attributes Improve Population Alignment in Large Language Models
- Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference
- Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
- Large Language Models Reflect the Ideology of their Creators
- Learning Pluralistic User Preferences through Reinforcement Learning Fine-tuned Summaries
Original note title
more than half of LLM personas fail to express their assigned World Values Survey profiles from the outset — drift is the smaller failure