Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference
As large language models are increasingly deployed to simulate diverse human characters, ensuring persona fidelity, defined as the extent to which an agent’s behavior consistently reflects the psychological and stylistic characteristics of a target persona, has become a critical requirement. However, existing evaluation paradigms primarily rely on either holistic LLM-based judges, which are prone to “holistic appraisal hallucination”, or static psychometric inventories, which fail to capture the context-dependent fidelity required in dynamic dialogue. To address these limitations, we propose PRISM (Persona Reasoning with Inverse SFL-based Modeling), a psycholinguistically grounded framework that reformulates persona fidelity evaluation as a structured inverse inference task. Inspired by Systemic Functional Linguistics (SFL), PRISM decomposes persona fidelity into three functional dimensions: Task Framing, Interpersonal Stance, and Linguistic Style. It estimates dimension-specific evidence over a persona-conditioned label space and aggregates these signals into an interpretable and auditable evaluation process.
Introduction. Recent advances in Large Language Models (LLMs) have enabled increasingly sophisticated role-playing agents that can simulate diverse personas and social identities (Tu et al., 2024; Li et al., 2025b). As these agents are increasingly deployed in immersive and interactive environments, ensuring their consistency with assigned characters has emerged as a crucial desideratum (Ji et al., 2025; Wang et al., 2024a; Bhandari et al., 2025). A compelling role-playing agent should not only generate coherent responses that remain consistent with persona-related knowledge, but also maintain stable and recognizable personality traits and behavioral styles throughout interaction (Wu et al., 2025a; Li et al., 2026a). This requirement is commonly referred to as persona fidelity: the extent to which a model’s behavior consistently reflects the psychological and stylistic characteristics of a target persona (Shin et al., 2025; Wang et al., 2024c). Despite its importance, reliably evaluating persona fidelity remains a significant challenge (Jiang et al., 2024; Yoon et al., 2024; Ji et al., 2025).
Discussion / Conclusion. We introduce PRISM, a persona-conditioned inverse structured evaluation framework for persona fidelity assessment. Grounded in Systemic Functional Linguistics, PRISM models personarelevant behavior along three functional dimensions: task framing, interpersonal stance and linguistic style. Across three benchmarks, PRISM consistently improves over direct LLM-as-a-judge baselines, especially on the harder benchmark and under stricter group-level metrics. These results suggest that persona fidelity is inherently multidimensional and is more reliably assessed through structured dimension-level evidence than through a single holistic judgement. While PRISM provides a structured and interpretable framework for persona fidelity evaluation, several limitations remain. Theoretical Scope of Functional Dimensions. PRISM adopts a psycholinguistic perspective grounded in Systemic Functional Linguistics (SFL) to organize persona-relevant behaviors into three dimensions: task framing, interpersonal stance and linguistic style.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can conversational AI maintain consistent personas across conversations?- How does persona consistency affect coherence in simulated dialogue?
- Do synthetic personas maintain consistency across multiple conversations?
- What are the three distinct types of persona drift in dialogue systems?
- What training objectives would actually improve persona consistency at scale?
- Can offline RL scale persona consistency across multi-turn conversations?
- How can training methods enforce persona consistency without supervised learning penalizing it?
- Can persona consistency coexist with relevant dialogue in personalized conversation?
- How does distractor persona selection affect consistency enforcement in dialogue?
- Why is persona consistency a pragmatic property rather than semantic?
- How does tree-structured persona maintenance prevent character drift in long conversations?
- Why does static persona definition fail to capture natural variation?
- Does persona assignment alone produce repetitive dialogue without situational grounding?
- Can persona prompts reliably transfer across different question domains?
- How do persona consistency and contextual relevance trade off in personalized dialogue systems?
- At what scale does persona distortion become a threat to public discourse?
- What makes extended personal narratives more effective than attribute lists for personas?
- Does linguistic style or content richness matter more for persona authenticity?
- How much does interview richness matter compared to model capability for persona accuracy?
- How should persona prompts be used if not for accuracy?
- Do individual persona simulations work?