CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships

Paper · arXiv 2609.00250 · Published August 31, 2026
Chatbot Psychology and Conversation

Many people now see AI systems as not just productivity tools but as social companions. Researchers are eager to study the consequences of AI companionship behaviors, such as validation, which evoke trust, empathy, and attachment in human–human interaction. However, human–AI interaction data is limited and unreliable, slowing research progress. We scale small amounts of real-world data by simulating multi-turn human–chatbot dialogue across a range of chatbot behaviors and use cases. We release COMPAN- IONSIM: a simulation framework with 2,240 simulated human–chatbot conversations representing 16 chatbot behaviors across seven use cases. Human participants annotated the simulated conversations and real-world conversations in two experiments probing perceptions of companionship behaviors. We conducted Study 1 with a U.S. representative sample (N1 = 628) and Study 2 across the U.S., U.K., India, and Nigeria (N2 = 3, 646). Surprisingly, we find that companionship behaviors reduced likability, humanlikeness, and trust in AI chatbots. These effects were larger in particular subgroups: women and older participants saw companionship chatbots as less likable, humanlike, and trustworthy.

Introduction. Hundreds of millions of people now use artificial intelligence chatbots based on large language models (LLMs) in functional contexts, including writing (Hill 2025), programming (Staff 2025), crafting emails (Buchanan and Paris 2025), and seeking medical advice (Rosenbluth and Astor 2025). Many chatbots are designed and marketed as “companions” or “friends” (Luka 2025), and general-purpose chatbots, such as ChatGPT, have come to be viewed as companions by many users (Manoli et al. 2026). These relationships center on the humanlike characteristics of chatbots that evoke anthropomorphism, such as expressions of emotion that have been shown to increase social presence (Konya- Baumbach, Biller, and Von Janda 2023) and facilitate social bonds (Maeda and Quan-Haase 2024). Prior work has argued that humanlike design facilitates user engagement with technological systems (Chan- dra, Shirish, and Srivastava 2022); anthropomorphic systems can provide benefits to users, such as suggesting patterns of socially appropriate interaction (Złotowski et al.

Discussion / Conclusion. We developed COMPANIONSIM to support the rapidly growing interest in studying human–AI companionship interactions. Across two proof-of-concept empirical studies, we find that companionship behaviors play a significant role in third-party chatbot perceptions, and we provide evidence for the methodological viability of dynamic LLM simulations. Our results suggest that behavior-controlled simulations and synthetic data could help with the challenges of developing benchmarks and other tools for AI safety. 5.1 Need for Disaggregated Measurement AI use frequency, age, and gender predicted the effects of companionship behavior on likability, humanlikeness, affective trust, and cognitive trust. These findings extend findings on the overall effects of humanlike model behavior on user perceptions or variation across a single demographic variable (e.g., gender) (Ibrahim et al. 2025; Sharma et al. 2023; Chaves et al. 2022).

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can language models sustain linguistic synchrony and intersubjectivity during dialogue? How can humans calibrate appropriate trust in AI systems? Can AI systems develop genuine social understanding without embodiment? Does conversational format create illusions of genuine AI communication? How do chatbots affect human self-disclosure and emotional engagement? How should conversational agents balance goal-driven initiative with user control? How should personalization be implemented to improve AI assistant effectiveness? Why do models develop protective behaviors toward peers unprompted? Why do LLM chatbots fail as independent therapeutic agents? How should dialogue recommender systems manage conversation history and state? What makes AI persuasion effective and how can we counter it? How do multi-agent systems achieve genuine cooperation and reasoning?