Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation

Paper · arXiv 2609.25010 · Published July 27, 2026
Personas and Personality

Marketers increasingly use large language models (LLMs) as “synthetic personas” to predict how an audience will react to a piece of copy before it ships, encouraged by evidence that profile-conditioned LLMs mimic human samples. But is that prediction actually valid against real behaviour—and does the persona machinery help? We present a sim-to-real validity study using the Upworthy Research Archive—thousands of headline A/B tests on shared real traffic, with measured click-through—as held-out ground truth. We compare a ten-persona panel, grounded in the real audience’s demographics, against a no-persona zero-shot baseline that simply asks the model how likely a typical reader is to click. Two findings stand out. First, ground-truth reliability is the binding constraint: most A/B tests have no statistically distinguishable winner, so validity can only be measured on the reliable subset (n = 399). Second, and counter to the persona-simulation premise, persona conditioning degrades predictive validity: the no-persona baseline ranks variants markedly better (Kendall τ = 0.361, a medium effect; top-1 accuracy 49.2%) than the persona panel (τ = 0.084; top-1 34.6%), with non-overlapping confidence intervals.

Introduction. Before launching a campaign, marketers want to know which version of a message will land. A fast- growing practice replaces (or precedes) live A/B testing with synthetic audience simulation: a large language model (LLM) is conditioned on a set of audience “personas” and asked to react to each candidate message, and the message its personas prefer is shipped. The appeal is obvious—instant, cheap feedback—and it is encouraged by findings that profile-conditioned LLMs can reproduce aspects of real human samples, an idea sometimes called “silicon sampling” [1]. The appeal, however, outruns the evidence. The premise that a synthetic persona predicts how a real audience behaves is rarely tested against real outcomes, because doing so requires ground truth that pairs concrete copy with measured audience response.

Discussion / Conclusion. Why persona conditioning hurts. The base model, asked directly, holds a usable population-level prior on what gets clicked— plausibly because clickability patterns are abundant in its pretraining data. Conditioning on a specific persona (“you are a 68-yearold retiree. . . ”) reframes the task as first-person roleplay, which substitutes an idiosyncratic, stereotyped guess for that prior; averaging across a ten-persona panel does not recover it, because each response is biased rather than merely noisy —a caricature effect documented for LLM persona simulations [2], where persona conditioning can surface implicit bias and degrade performance on objective tasks [7]. Seen another way, our no-persona baseline is itself a single aggregate persona (a “typical reader”); the finding is then that disaggregating the audience into a demographic panel hurts an aggregate prediction task— We tested whether persona-based copy simulation predicts how a real audience ranks marketing copy, using the Upworthy A/B-test archive as held-out ground truth, and whether the persona machinery helps at all. Two findings stand out.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can persona representations reduce language model variance and improve task accuracy? How should personalization be implemented to improve AI assistant effectiveness? How can LLM user simulators model realistic goal-driven conversation? Why do persona-level simulations fail to predict individual preferences accurately? How do evaluation biases undermine LLM quality assessment systems? Do language models develop causal world models or rely on statistical patterns? What prevents language models from reliably adopting diverse personas? How can recommendation systems balance personalization with stability and coverage? Is model self-awareness based on genuine introspection or pattern matching? Can AI-generated outputs constitute genuine knowledge or valid claims? Why should disagreement be treated as signal in collaborative reasoning? Do accurate-looking LLM outputs hide structural failures in learning and reasoning? How can conversational AI maintain consistent personas across conversations? Can AI systems develop genuine social understanding without embodiment?