Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation
Marketers increasingly use large language models (LLMs) as “synthetic personas” to predict how an audience will react to a piece of copy before it ships, encouraged by evidence that profile-conditioned LLMs mimic human samples. But is that prediction actually valid against real behaviour—and does the persona machinery help? We present a sim-to-real validity study using the Upworthy Research Archive—thousands of headline A/B tests on shared real traffic, with measured click-through—as held-out ground truth. We compare a ten-persona panel, grounded in the real audience’s demographics, against a no-persona zero-shot baseline that simply asks the model how likely a typical reader is to click. Two findings stand out. First, ground-truth reliability is the binding constraint: most A/B tests have no statistically distinguishable winner, so validity can only be measured on the reliable subset (n = 399). Second, and counter to the persona-simulation premise, persona conditioning degrades predictive validity: the no-persona baseline ranks variants markedly better (Kendall τ = 0.361, a medium effect; top-1 accuracy 49.2%) than the persona panel (τ = 0.084; top-1 34.6%), with non-overlapping confidence intervals.
Introduction. Before launching a campaign, marketers want to know which version of a message will land. A fast- growing practice replaces (or precedes) live A/B testing with synthetic audience simulation: a large language model (LLM) is conditioned on a set of audience “personas” and asked to react to each candidate message, and the message its personas prefer is shipped. The appeal is obvious—instant, cheap feedback—and it is encouraged by findings that profile-conditioned LLMs can reproduce aspects of real human samples, an idea sometimes called “silicon sampling” [1]. The appeal, however, outruns the evidence. The premise that a synthetic persona predicts how a real audience behaves is rarely tested against real outcomes, because doing so requires ground truth that pairs concrete copy with measured audience response.
Discussion / Conclusion. Why persona conditioning hurts. The base model, asked directly, holds a usable population-level prior on what gets clicked— plausibly because clickability patterns are abundant in its pretraining data. Conditioning on a specific persona (“you are a 68-yearold retiree. . . ”) reframes the task as first-person roleplay, which substitutes an idiosyncratic, stereotyped guess for that prior; averaging across a ten-persona panel does not recover it, because each response is biased rather than merely noisy —a caricature effect documented for LLM persona simulations [2], where persona conditioning can surface implicit bias and degrade performance on objective tasks [7]. Seen another way, our no-persona baseline is itself a single aggregate persona (a “typical reader”); the finding is then that disaggregating the audience into a demographic panel hurts an aggregate prediction task— We tested whether persona-based copy simulation predicts how a real audience ranks marketing copy, using the Upworthy A/B-test archive as held-out ground truth, and whether the persona machinery helps at all. Two findings stand out.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can persona representations reduce language model variance and improve task accuracy?- Do individual persona simulations work?
- How do LLM personas compare to demographic targeting?
- Why do short interviews outperform demographic labels for persona simulation?
- Why do individual persona simulations succeed when population-level representation fails?
- Can agent-based simulators replace real-user A/B testing for studying recommendation system harms?
- How do LLM user simulators fail to represent authentic user behavior distributions?
- Can structured empathy measurement frameworks predict persona effectiveness?
- What makes personas in multi-agent systems actually contribute meaningful domain depth?
- Does adding survey data to interviews improve agent accuracy further?
- How much does omniscient evaluation overstate real-world simulation fidelity?
- How does support coverage relate to systematic biases in persona simulation?
- What distribution patterns appear across different theory-of-mind datasets?
- What role does authentic self-expression play in building accurate personality models?