CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
Many people now see AI systems as not just productivity tools but as social companions. Researchers are eager to study the consequences of AI companionship behaviors, such as validation, which evoke trust, empathy, and attachment in human–human interaction. However, human–AI interaction data is limited and unreliable, slowing research progress. We scale small amounts of real-world data by simulating multi-turn human–chatbot dialogue across a range of chatbot behaviors and use cases. We release COMPAN- IONSIM: a simulation framework with 2,240 simulated human–chatbot conversations representing 16 chatbot behaviors across seven use cases. Human participants annotated the simulated conversations and real-world conversations in two experiments probing perceptions of companionship behaviors. We conducted Study 1 with a U.S. representative sample (N1 = 628) and Study 2 across the U.S., U.K., India, and Nigeria (N2 = 3, 646). Surprisingly, we find that companionship behaviors reduced likability, humanlikeness, and trust in AI chatbots. These effects were larger in particular subgroups: women and older participants saw companionship chatbots as less likable, humanlike, and trustworthy.
Introduction. Hundreds of millions of people now use artificial intelligence chatbots based on large language models (LLMs) in functional contexts, including writing (Hill 2025), programming (Staff 2025), crafting emails (Buchanan and Paris 2025), and seeking medical advice (Rosenbluth and Astor 2025). Many chatbots are designed and marketed as “companions” or “friends” (Luka 2025), and general-purpose chatbots, such as ChatGPT, have come to be viewed as companions by many users (Manoli et al. 2026). These relationships center on the humanlike characteristics of chatbots that evoke anthropomorphism, such as expressions of emotion that have been shown to increase social presence (Konya- Baumbach, Biller, and Von Janda 2023) and facilitate social bonds (Maeda and Quan-Haase 2024). Prior work has argued that humanlike design facilitates user engagement with technological systems (Chan- dra, Shirish, and Srivastava 2022); anthropomorphic systems can provide benefits to users, such as suggesting patterns of socially appropriate interaction (Złotowski et al.
Discussion / Conclusion. We developed COMPANIONSIM to support the rapidly growing interest in studying human–AI companionship interactions. Across two proof-of-concept empirical studies, we find that companionship behaviors play a significant role in third-party chatbot perceptions, and we provide evidence for the methodological viability of dynamic LLM simulations. Our results suggest that behavior-controlled simulations and synthetic data could help with the challenges of developing benchmarks and other tools for AI safety. 5.1 Need for Disaggregated Measurement AI use frequency, age, and gender predicted the effects of companionship behavior on likability, humanlikeness, affective trust, and cognitive trust. These findings extend findings on the overall effects of humanlike model behavior on user perceptions or variation across a single demographic variable (e.g., gender) (Ibrahim et al. 2025; Sharma et al. 2023; Chaves et al. 2022).
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can language models sustain linguistic synchrony and intersubjectivity during dialogue? How can humans calibrate appropriate trust in AI systems?- How does outcome feedback change beliefs about AI versus human partner reliability?
- Can validation procedures interrupt an AI's relationship-maintenance logic?
- How do Heersmink's integration dimensions explain why chatbots feel more trustworthy than other tools?
- How does consciousness attribution drive emotional dependence on chatbots?
- How does emotional dependence on chatbots affect user wellbeing?
- How do user expectations change as chatbots remember more interactions?
- How does the expectation ratchet affect long-term chatbot satisfaction?
- What temporal design dimensions characterize different chatbot relationship types?
- Why do persistent chatbot companions face novelty decay that ad-hoc supporters avoid?
- Can transparency about AI limitations reduce the seductiveness of chatbots as quasi-Others?
- Does chatbot interaction reduce authentic personal expression in dialogue?
- How does understanding persistent journeys intensify both trust and privacy concerns?
- Does personalization help or hurt persistent companion chatbots?
- How do intrinsic motivation mechanisms differ between social proactivity and personalization?