SYNTHESIS NOTE
Topics›Personas Personality›this note

Why do persona prompts show such mixed results for surveys?

Persona prompting produces inconsistent outcomes when simulating survey responses. This explores whether variation in how humans answer specific questions might explain when the technique works and when it fails.

Synthesis note · 2026-09-25 · sourced from Personas Personality

The paper starts from an unresolved pattern: persona prompting, meaning short textual descriptions of socio-demographics, attitudes or behaviors placed in the prompt, has produced "mixed and partly conflicting results" when LLMs are asked to predict survey-panel responses, "without clear patterns of success and failure." The few consistent findings are that the choice of persona attributes matters and that more attributes do not necessarily mean better performance. The authors propose that "observed human response variation of a survey question" is "a potential explanation" for the mixed record. Their discussion carries this to a practical claim: simulation "will generally" work best for questions with high human response variation, that is, questions that are "highly contested."

The reasoning is a two-step framework for practitioners. First, measures of human response variation triage which survey questions (and potentially tasks and topics) are good candidates for persona simulation at all. Second, for the questions that pass, attributes should be chosen with "approaches informed by existing survey data for the same or a similar question," because those approaches identify the attributes most important for persona creation and give the best alignment between predicted and actual responses. The paper also compares several attribute-selection methods, which is its answer to the open question of how to choose among them. In this framing, whether a persona helps depends on a property of the question, not only on the model or the prompt.

This sits beside several existing notes as a qualifier rather than a rival. Does conditioning LLMs on personal profiles improve prediction? shows the individuation lever coming up empty at the level of single participants; this paper's target is alignment with human survey responses, and it suggests the payoff varies by question, which is compatible with weak per-person gains. Why do LLM persona prompts produce inconsistent outputs across runs? draws on annotation work, and the introduction here notes that the limited evidence on sources of mixed results comes "primarily from subjective annotation tasks"; this paper moves the question to survey panels. Why do LLMs give unrealistic survey responses? locates survey-simulation failure in how responses are elicited, whereas this paper locates part of it in which questions are asked. And How do we generate realistic personas at population scale? names unresolved persona-attribute choice as a gap; the survey-data-informed selection step is one concrete answer to it.

The excerpt does not establish the details. It does not name the models, surveys or attribute-selection methods, say how response variation was measured, report effect sizes, or say how much of the mixed performance the variation account explains. The abstract offers the account as "a potential explanation," and the conclusion says the results "suggest" the framework, so the strength is a supported hypothesis with a practical rule of thumb, not a settled mechanism. The implication that follows: before trusting a persona simulation, check how contested the question is among real respondents, and do not assume adding attributes helps. Where humans largely agree, the excerpt gives no reason to expect persona conditioning to add value.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What makes personas effective for predicting individual preferences and behavior? How can conversational agents maintain consistent personas across multi-turn dialogue?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 54 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

human response variation may explain mixed persona prompting results — simulation is expected to work best on contested survey questions