Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA’s output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles.
Introduction. Role-playing agents (RPAs) provide interactive storytelling, companionship, and emotionally engaging conversation. Their value often unfolds over multiple turns, and deployed systems already serve millions of users (Irvine et al., 2023). As these systems enter large-scale real-world use, reliable evaluation becomes essential for measuring capability, comparing systems, and guiding further improvement. Unlike general-purpose task assistants, the object being evaluated is not an isolated response but a dialogue jointly constructed by the RPA and its user. Such dialogues rarely have a uniquely correct answer, and conventional automatic metrics correlate poorly with human judgments in open-ended dialogue (Liu et al., 2016). Quality must instead be established through measures of character fidelity, narrative development, and user experience.
Discussion / Conclusion. Fixed-history evaluation places an RPA in a trajectory that it did not help construct, thereby mixing its own capability with the influence of the external history. User-independent scoring further compresses genuine individual differences in satisfaction into a single standard. Palate uses simulators trained from real user histories to generate free multi-turn dialogues and evaluates candidates along personalized, generic turn-level, and whole-session tracks. It further decomposes one person into a behavior policy learned from real dialogue and an individual utility supervised by experience annotations, changing the basic unit of role-play evaluation from an isolated RPA to a user–RPA pair. This introduces a user-perspective satisfaction reference while retaining general quality evaluation. Across 16 candidates, advantages on the three tracks do not coincide, and the five users do not share a single best candidate. The central output of Palate is therefore not another scalar leaderboard, but an interactive evaluation profile that locates cross-track capability mismatches and per-user differences.
Lines of inquiry this paper opens 8
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do interface design choices shape consciousness attribution?- What measurable harms occur when users interact with AI as if it were conscious?
- How does the philosophical distinction between simulation and realization affect liability?