INQUIRING LINE

Give an AI a persona with an age, gender and country: does it actually adopt that person's values?

How do demographic details affect whether personas hold assigned value profiles?

This explores whether adding demographic details (age, gender, country and similar) to a persona helps or hurts its ability to actually express the value profile it was assigned.


This explores whether adding demographic details to a persona helps it express the values it was assigned. The corpus has no study that varies demographic detail and then measures value-holding, so it can't answer directly. What it does have is a striking finding about where value-holding breaks, plus several results suggesting demographic detail is a weak lever.

The key finding is about timing. Across 4,000 conversations and three models, more than half of personas never instantiated their assigned World Values Survey profile at all, while only 2 to 7 percent drifted away from it later (Do LLM personas actually embody their assigned cultural values?). The failure happens at the start, and the conversations still read as plausible. So the question isn't which details keep a persona from wandering off. It's whether the details get the values in to begin with, and for most personas they don't.

Other work points the same way. Conditioning models on personal profiles did not improve predictions for specific individuals across more than 200,000 participants (Does conditioning LLMs on personal profiles improve prediction?). A ten-persona demographic panel ranked headlines worse than a single 'typical reader' prompt with no persona, which suggests demographic conditioning adds bias rather than audience insight (Do demographic personas help models rank headlines better?). Persona prompts also change what the model says without changing the bias underneath, since between-group gaps persisted unchanged (Can persona prompts actually reduce bias in language models?). A demographic label seems to steer the surface of the output, not the model's value structure.

There are some hints at why. A list of demographics gives the model marginal facts, such as age, gender and country, but values depend on how those combine. LLM persona generation runs on heuristics that can't recover true joint distributions from marginal data (How do we generate realistic personas at population scale?). When models have thin information they also fall back on stereotype defaults. That result comes from the reverse task, inferring demographics from sparse social media accounts (Can LLMs predict demographics from social media usernames alone?). A separate line of work argues that post-training installs stable dispositional profiles that persist under pressure (Are RLHF personas performed characters or realized dispositions?). It is tempting to read that as the model's own disposition overriding an assigned one, but none of these notes tests that.

The contrast that does show up is grounding. Personas built from real behavioral data predicted A/B test directions at 75 to 90 percent accuracy (Can behavior-based personas predict A/B test outcomes?). Demographic labels alone did much worse in the headline test above. Near-matches can also be dangerous: in personalization, profiles that are almost right produced the worst errors, because the model applies the wrong preferences with confidence (Why do similar user profiles produce worse personalization errors?). A nearly right demographic sketch could plausibly do the same to values. The corpus doesn't test that, so a controlled experiment that adds or removes demographic fields and scores against the assigned value profile would be a genuine gap.


Sources 0 notes