Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

Paper · arXiv 2608.06485 · Published August 6, 2026
Personas and Personality

Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event–trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples.

Introduction. Personality-conditioned LLM agents (PC-Agents) have become a foundational primitive across a growing list of applications (Chen et al., 2026). They power emotional companionship and mentalhealth support chatbots (Hu et al., 2026), populate social simulations (Mou et al., 2026) for behavioral research (Park et al., 2026; Larooij and Törnberg, 2025), and serve as long-horizon role-playing engines for interactive fiction, games, and digital tutoring (Park et al., 2023; Chen et al., 2024). Multi-session systems such as AnnaAgent already couple evolving emotional and cognitive states with persistent memory in psychological counselling (Wang et al., 2025). Across these settings, lifelong agents must maintain a coherent persona across extended interactions, long wall-clock durations, and open-ended user-driven narrative arcs. A foundational question for lifelong PC-Agents is how their personality evolves relative to humans. In human psychology, personality traits, though enduring, can change due to life events.

Discussion / Conclusion. This paper explored whether PC-Agents exhibit psychologically plausible personality evolution after major life events. We measured their Big Five profiles before and after each event and used the resulting trajectories to construct BFI-Adapt. Current models can move, but their movement is weakly eventspecific, poorly calibrated in magnitude, and compressed across demographic and individual variation. Present PC-Agents therefore approximate a generic pattern of change more readily than the event- and person-specific structure of human personality development. Event-conditioned trajectories exceed retest noise, preserve their event–trait structure under independent paraphrases, exhibit model-dependent convergence with scenario-based decisions, and remain detectable after unrelated dialogue. These complementary checks validate the trajectory analysis and its central conclusions.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems balance emotional competence with factual reliability? What prevents language models from reliably adopting diverse personas? Can AI systems develop genuine social understanding without embodiment? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? How can conversational AI maintain consistent personas across conversations? What makes AI persuasion effective and how can we counter it? Is model self-awareness based on genuine introspection or pattern matching? What limits mechanistic interpretability's ability to characterize models? Do language model representations contain causally steerable task-specific features? Why do LLM chatbots fail as independent therapeutic agents? What structural biases does transformer attention create in language model outputs? Does RLHF training sacrifice accuracy and grounding for user agreement? Do language models develop causal world models or rely on statistical patterns?