Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character descriptions that fail to sustain coherent character behavior across extended interactions. We introduce Deep Persona, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents. Governed by the principles of scripted determinism and bounded agency, the architecture restricts the model to a reactive engine guided by a structured internal script. We further propose a referencefree evaluation framework that benchmarks dialogue naturalness against empirical human distributions using established psychological clinical instruments and adversarial stress-tests. Empirical evaluation reveals that while LLMs achieve high pragmatic fluency, they exhibit systematic limitations in emotional expression and joint attention. In addition, we present a case study of two Deep Personas and evaluate them using the proposed framework, demonstrating that structured personas can produce interactions that more closely align with human conversational behavior.
Introduction. LLMs are increasingly used as interactive agents capable of simulating real people across a wide range of domains (Tseng et al., 2024). These systems support applications such as conversational assistants, educational tools, entertainment platforms, and training environments, where models are expected to adopt specific identities and engage in multi-turn interactions. Recent work has demonstrated the ability of LLMs to perform role-playing in conversational settings (Tao et al., 2024; Wang et al., 2024a; Zhou et al., 2025). In particular, LLM-based agents are increasingly explored in mental health and clinical training, where simulated interactions support the education, evaluation, and skill development of therapists (Lawrence et al., 2024; Hua et al., 2025; Elyoseph et al., 2026). These applications place strong demands on the realism, consistency, and stability of simulated personas. Despite this progress, current approaches to persona modeling remain fundamentally limited.
Discussion / Conclusion. • A central design goal of our framework is to maintain consistent persona behavior, even under adversarial conditions. However, this objective may conflict with safety requirements, particularly when interactions involve harmful, aggressive, or sensitive content. Ensuring robust system-level safeguards is therefore essential to prevent inappropriate or unsafe outputs. • The ability of the system to generate highly human-like interactions also introduces the risk of misuse. In uncontrolled settings, such capabilities could be used to deceive users or obscure the artificial nature of the agent.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can persona representations reduce language model variance and improve task accuracy?- Do individual persona simulations work?
- Why do language models successfully simulate political perspectives and social personas?
- How do LLM personas compare to demographic targeting?
- At what scale does persona distortion become a threat to public discourse?
- How does behavioral stickiness distinguish realized from pretended personas?
- Can one model instance host multiple realized personas simultaneously?
- How does persona consistency affect coherence in simulated dialogue?
- Can fine-tuning or RLHF alone solve the persona distortion problem?
- Do synthetic personas maintain consistency across multiple conversations?
- Can controllable latent variables in simulators ground them to realistic conversation?
- What role does user contribution play in constituting the interlocutor?
- How does non-human origin of personas affect team willingness to critique them?
- Can structured empathy measurement frameworks predict persona effectiveness?
- Can synthetic personas achieve emotional connection with creators?