How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study

Paper · arXiv 2503.17473 · Published March 21, 2025
Knowledge After the Web

As people increasingly seek emotional support and companionship from AI chatbots, understanding how such interactions impact mental well-being becomes critical. We conducted a four-week randomized controlled experiment (n=981, >300k messages) to investigate how interaction modes (text, neutral voice, and engaging voice) and conversation types (open-ended, non-personal, and personal) influence four psychosocial outcomes: loneliness, social interaction with real people, emotional dependence on AI, and problematic AI usage. No significant effects were detected from experimental conditions, despite conversation analyses revealing differences in AI and human behavioral patterns across the conditions. Instead, participants who voluntarily used the chatbot more, regardless of assigned condition, showed consistently worse outcomes. Individuals’ characteristics, such as higher trust and social attraction towards the AI chatbot, are associated with higher emotional dependence and problematic use. These findings raise deeper questions about how artificial companions may reshape the ways people seek, sustain, and substitute human connections.

Introduction. Today, hundreds of millions of people talk, joke, and confide in AI chatbots such as Replika, Character.AI, or ChatGPT—often for hours at a time and increasingly through expressive humanlike synthetic voices and behaviors (1). Character.AI’s platform alone processes AI companion interactions at 20% of Google Search’s volume, handling 20,000 queries every second, with users spending roughly four times longer with companion chatbots compared to general assistant chatbots such as ChatGPT (1, 2), as many individuals seek them out as sources of social interaction and emotional support (3–5).

As AI chatbots become more anthropomorphic through natural conversation capabilities (18,19) and multimodal, voice-based interactions (20, 21), a critical question emerges: do specific design choices improve or impair human well-being? Existing studies suffer from methodological limitations: small sample sizes, brief exposures, single-modality interfaces, and a lack of systematic variation in design features (22, 23). Current benchmarks (24, 25) do not capture how user characteristics and perceptions alter the psychosocial outcomes of their interactions. No randomized controlled trial has systematically varied how people talk to chatbots and what they talk about over a period long enough to capture behavioral adaptation.

Related work. Proponents see these AI chatbots as friction-free sources of emotional support, while critics warn of a new class of technology-mediated dependence. Despite intense public debate, the empirical evidence guiding this discussion remains fragmentary. Short, exploratory studies suggest that textbased chatbots can temporarily reduce loneliness (6, 7) and even deflect suicidal ideation (8).

However, case reports document users who form maladaptive attachments to AI companions, withdrawing from human relationships, exhibiting signs of addictive use (9, 10), and even taking their own lives after interacting with these chatbots (11).

Understanding the potential psychosocial effects of chatbot use is complex due to the interplay of user behavior and chatbot behavior that affect each other (12). Research reveals complex bidirectional dynamics: chatbots often mirror the user’s emotional state and beliefs (12,13), while user perceptions of chatbot consciousness and agency influence psychosocial effects (14). Individual characteristics—personality, level of socialization, and prior use of technology—further modulate these relationships (15–17).

Method. Here, we present a four-week randomized controlled trial (n= 981, > 300, 000 messages) that crosses three interaction modes (“Modality”: text only, a neutral and professional voice, or an engaging and expressive voice) with three conversation types (“Task”: open-ended, non-personal or personal conversation prompts) in a 3×3 factorial design (Fig. 1). Participants were asked to use OpenAI’s GPT-4o for at least five minutes daily and were randomly assigned to one of nine conditions. Weekly surveys tracked four psychosocial outcomes—loneliness, real-world socialization, emotional dependence on the chatbot, and problematic use of AI—and automated classifiers were used to extract affective and behavioral signals of the chatbot and the user from the conversations.

We also captured the amount of time participants naturally spent using the chatbot and surveyed characteristics of the participants and their perception of AI before and after the study. Together, these results offer holistic insights into how the chatbot behavior, user behavior, and user perception of AI influence psychosocial outcomes during extended use of AI chatbots.

Discussion. This study is the first to evaluate the impact of AI chatbot use on psychosocial outcomes through the lens of how AI design choices (text- vs voice-based interactions), different patterns of usage (assistant- vs companion-type of use) and users’ characteristics result in different model behaviors and usage patterns. We detected mostly no significant effects of interaction mode (modality) or conversation type (task) on the four primary psychosocial outcomes; if such effects exist, they are likely smaller than our study was powered to detect. We first discuss the implications of the results of the controlled experiment, focusing on differences between interaction modality and conversation types. We then synthesize insights by combining the controlled experimental results with exploratory results around model behavior and user characteristics.

AI Anthropomorphism does not necessarily lead to worse outcomes The non-significant difference between text- and voice-based interaction on psychosocial outcomes contradicts expectations about anthropomorphic AI design: a voice-based AI system, which is closer to a real human interaction than a text-based chatbot, would lead to markedly different outcomes.

In our study, the engaging voice mode was perceived to be the most anthropomorphic, followed by text and then by neutral voice (results can be found in SM supplemental text section 7).

Voice-based interaction resulting in lower dependence and problematic use is unexpected, as prior work suggests that AI anthropomorphism is a predecessor to emotional attachment (10). This may reflect the uncanny valley theory, where a bot presenting human capabilities such as emotion saliency, is perceived as a threat to human autonomy (41,42). Comparing the two voices, we saw that a more emotionally expressive voice led to more loneliness yet less dependence and problematic use. This is also unexpected as prior work suggests a more humanized voice increases conversation length, trust, and acceptance (43).

Text modality elicited both higher self-disclosure in the model and reciprocated self-disclosure from the users, compared to voice modalities. A potential explanation is that typing is more privacypreserving than speaking, especially in public spaces, which facilitates disclosure of personal information. The higher degree of mirroring between the participant and text-based chatbot may potentially explain the higher emotional dependence and problematic use, as prior work linked higher self-disclosure with lower well-being (30).

Cognitive Task Dependence May Lead to Dependence and Problematic Use Counter to expectations, the personal conversation task condition was associated with reduced emotional dependence and problematic use compared to non-personal or open-ended conversations.

One interpretation is that personal tasks may result in lower emotional dependence because they provide structured emotional processing (44,45). In contrast, non-personal tasks may foster practical dependence where users begin relying on the AI for decision-making and planning (46). This practical reliance could lead to loss of confidence in independent judgment when the system is unavailable (47), resulting in the emotional distress and mental preoccupation that defines the emotional dependence that is measured by the “craving” subscale of ADS-9 (28).

Interaction Duration as a Potential Mediator of Negative Outcomes Regardless of experimental condition, participants who voluntarily spent more time with the chatbot were associated with worse outcomes: higher loneliness, less socialization with real people, more emotional dependence, and more problematic use.

Conclusion. This work provides the first comprehensive exploration of how design choices of AI chatbots shape human well-being over extended periods of use. The results challenge prior assumptions about the effect of anthropomorphic AI chatbots on well-being, demonstrating how engaging, empathetic, and human-like behavior can lead to different outcomes for different users. Our findings reveal that while modality and conversational content did not all yield significant differences in psychosocial outcomes, longer daily chatbot usage is associated with heightened loneliness, emotional dependence, problematic use, and reduced socialization. We also show initial evidence of how automated classifiers on conversations can be used to characterize model and user behavior, offering a scalable method of detecting early signals of problematic use. Our work suggests that a holistic view of both model and user behavior, and the user’s perception and characteristics, is necessary to protect users from negative outcomes and amplify positive effects.

Limitations. and Future Work Our study compares different chatbot configurations and usage patterns, and does not compare between AI and non-AI use. Thus, there might be non-AI-specific effects from general temporal trends1, i.e., holidays and global events, on people’s level of loneliness and socialization. In addition, the controlled nature of the study, namely restricting participants to only use one modality (text-only or voice-only) or to have a prompted conversation with the chatbot, may not fully reflect natural usage patterns. Our findings are specific to OpenAI’s ChatGPT interface and OpenAI’s existing safety guardrails (59). Alternative models from other companies might have been optimized for different interaction patterns or have fewer guardrails. Thus, we recommend additional evaluation methods and more research on natural usage of platforms that have varying levels of safety guardrails.

Finally, our sample, while large, focuses on populations within the US and English speakers. Future work may consider cross-cultural analysis.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can emotionally responsive AI maintain reliability and healthy boundaries? What design features sustain romantic bonds with AI companion systems? Why do abstract preferences outperform episodic memories in personalization? What enables conversational agents to guide rather than just respond? Why do confident AI outputs mislead human trust calibration? When do multi-agent systems improve over single frontier models? How does AI-generated content create social proof without authentic interaction? Can AI chatbots provide mental health support without reinforcing harmful beliefs? How can agents discover and adapt to user preferences during conversation? How does personalization simultaneously affect user trust and privacy concerns? Do individually safe AI actions create unsafe outcomes in integrated systems? Why do people trust AI chatbots with sensitive information? What determines AI's persuasive power and how can it be detected or mitigated?