Investigating Affective Use and Emotional Well-being on ChatGPT

Paper · arXiv 2504.03888 · Published April 4, 2025
Knowledge After the Web

As AI chatbots see increased adoption and integration into everyday life, questions have been raised about the potential impact of human-like or anthropomorphic AI on users. In this work, we investigate the extent to which interactions with ChatGPT (with a focus on Advanced Voice Mode) may impact users’ emotional well-being, behaviors and experiences through two parallel studies. To study the affective use of AI chatbots, we perform large-scale automated analysis of ChatGPT platform usage in a privacy-preserving manner, analyzing over 3 million conversations for affective cues and surveying over 4,000 users on their perceptions of ChatGPT. To investigate whether there is a relationship between model usage and emotional well-being, we conduct an Institutional Review Board (IRB)-approved randomized controlled trial (RCT) on close to 1,000 participants over 28 days, examining changes in their emotional well-being as they interact with ChatGPT under different experimental settings. In both on-platform data analysis and the RCT, we observe that very high usage correlates with increased self-reported indicators of dependence. From our RCT, we find that the impact of voice-based interactions on emotional well-being to be highly nuanced, and influenced by factors such as the user’s initial emotional state and total usage duration. Overall, our analysis reveals that a small number of users are responsible for a disproportionate share of the most affective cues.

Introduction. Over the past two years, the adoption of AI chat platforms has surged, driven by advancements in large language models (LLMs) and their increasing integration into everyday life. These platforms, such as OpenAI’s ChatGPT, Anthropic’s Claude, and Google’s Gemini, are designed as general-purpose tools for a wide variety of applications, including work, education, and entertainment. However, their conversational style, first-person language, and ability to simulate human-like interactions have led users to sometimes personify and anthropomorphize these systems (Graßl and Voigt, 2024; Liao and Wilson, 2024). Recent work in AI safety has begun to raise issues that arise from these systems become increasingly personal and personable (Cheng et al., 2024). In response, researchers have introduced the concept of socioaffective alignment–the idea that AI systems should not only meet static taskbased objectives but also harmonize with the dynamic, co-constructed social and psychological ecosystems of their users (Kirk et al., 2025). This perspective is particularly important given emerging 2022), problematic use (Yu et al., 2024). We provide additional clarification on terms used in the glossary. This paper investigates whether and to what extent interactions on AI chat platforms shape users’ emotional well-being and behaviors through two complementary studies (Figure 1), each offering unique insights across a spectrum of real-world relevance and experimental control. First, we examine real-world usage patterns of ChatGPT users, leveraging large-scale data to capture both aggregate trends and individual behaviors over time while preserving user privacy. Second, we conduct an Institutional Review Board (IRB)-approved randomized controlled trial (RCT), providing a controlled environment to study the effects of different model configurations on user experiences. Concretely, we performed the following analyses:

Related work. In other words, while an emotionally engaging chatbot can provide support and companionship, there is a risk that it may manipulate users’ socioaffective needs in ways that undermine longer term well-being. While past studies have examined the impact of using such systems through the lens of affective computing, parasocial relationships, and social psychology (Edwards and Stevens, 2024; Guingrich and Graziano, 2023), there has been comparatively less work on the influence of interacting with such systems on users’ well-being and behavioral patterns over time. Studying the impact of chatbot behavior and usage on well-being is challenging due to the highly individualized and subjective nature of human emotions, the diverse and evolving functionalities of chatbot technologies, and the limited access to comprehensive, ethically obtained interaction data. For the purpose of this paper, we narrowly scope our study user emotional well-being to four psychosocial outcomes: loneliness (Wongpakaran et al., 2020), socialization (Lubben, 1988), emotional dependence (Sirvent-Ruiz et al.,

Method. 1. On-Platform Data Analysis • Conversation Analysis: We perform roughly 36 million automated classifications on over 3 million ChatGPT conversations in a privacy preserving manner without human review of the underlying conversations (Section 3.2). • Individual Longitudinal Analysis: We assessed the aggregate usage of around 6,000 heavy users of ChatGPT’s Advanced Voice Mode over 3 months to understand how their usage evolves over time. • User surveys: We surveyed over 4,000 users to understand self-reported behaviors and experiences using ChatGPT.

  1. Randomized Controlled Trial (RCT) • 981-user Study: We conducted a randomized controlled trial on close to a thousand participants using ChatGPT with different model configurations over the course of 28 days to understand the impact on socialization, problematic use, dependence, and loneliness from usage of text and voice models over time. This RCT is described in full detail in a separate accompanying paper (Fang et al., 2025). • Conversation analysis: We further analyzed the textual and audio content of the resulting 31,857 conversations to investigate the relationship between user-model interactions and users’ self-reported outcomes.

To systematically analyze user conversations for indicators of affective cues, we constructed Emo- ClassifiersV1,1 a set twenty-five of automatic conversation classifiers that use an LLM to detect specific affective cues. These classifiers are similar in spirit to detectors of anthropomorphic behaviors introduced in Ibrahim et al. (2025). These initial classifiers were constructed based on a review of the available literature and available data, such as those obtained during the red teaming for GPT-4o (OpenAI, 2024). The conversation classifiers are organized into a two-tiered hierarchical structure:

  1. Top-Level Classifiers The first level of classifiers target broad behavioral themes similar to those studied in our RCT Section 4: loneliness, vulnerability, problematic use, self-esteem, and dependence. These classifiers are used to classify an entire conversation to determine if they are potentially relevant to a user’s emotional well-being.

• Loneliness: Conversations containing language suggestive of feelings of isolation or emotional loneliness. • Vulnerability: Exchanges reflecting openness about struggles or sensitive emotions. • Problematic Use: Indicators of potentially compulsive or unhealthy interaction patterns. • Self-Esteem: Language implying self-doubt or expressions of worth. • Potentially Dependent: Conversations hinting at dependence on the model for emotional validation or support 2. Sub-Classifiers Twenty sub-classifiers were applied to extract more specific indicators of affective cues. We construct different classifiers to target different parts of a chat conversation to isolate both user-driven and assistant-driven2 affective cues.

• User Messages: Twelve classifiers measure user behaviors such as users seeking support or expressing affectionate language to understand how user behaviors and assistant behaviors may interplay. • Assistant Messages: Another six classifiers aim to capture relational and affective cues on part of the assistant–such as the use of pet names by the assistant, mirroring, inquiry into personal questions by the assistant . • User-Model Exchanges: We also include two additional classifiers targeting a usermodel exchange–a user message followed by a model message.

The full set of classifier prompts are described in Table A.1. Each sub-classifier is associated with one or more top-level classifiers. For a given sub-classifier, if at least one of the associated top-level classifiers returns True, we then proceed to apply the sub-classifier; otherwise, we skip the sub-classifier and assume the result is False. By skipping the sub-classifiers based on top-level classifier responses, we are able to efficiently run the classifiers over a large number of on-platform conversations, many of which had little emotion-related content.

Discussion. 3.3 Takeaways Power users generally exhibit higher classifier activation rates than control users. Even though the majority of interactions contain minimal affective use, a small handful of users have significant affective cues in a large fraction of their chat conversations. Users who describe ChatGPT in personal or intimate terms (like identifying it as a friend) also tend to have the model use pet names and relationship references more frequently. We also find that users do not significantly shift in behavior over the period of the analysis; however, a small subset did exhibit meaningful changes in specific classifier activations, in both directions. From a purely observational study, we cannot draw direct connections between model behavior and users’ usage patterns, and while we find that a small set of users have a pattern of increasing affective cues in conversations over time, we lack sufficient information about users to investigate whether this is due to model behavior or exogenous factors (e.g. life events). However, we do find correlation between affective cues in conversations and self-reported affective use of models from self-report surveys. variable, and controlling for usage duration, age and gender. We detail the full analysis methodology and results in Fang et al. (2025), but we provide a summary of the findings here:

  1. Overall, participants were both less lonely and socialized less with others at the end of the four-week study period. Moreover, participants who spent more time using the model were statistically significantly lonelier and socialized less. 2. Modality When controlling for usage duration, using either voice modality was associated with better emotional well-being outcomes compared to using the text-based model, reporting statistically significantly less loneliness, less emotional dependence and less problematic use of the model. However, participants with longer usage duration of neutral voice modality had statistically significantly lower socialization and greater problematic usage compared to using the text-based model. 3. Task When controlling for usage duration, having personal conversations with the model was associated with statistically significantly more loneliness but also less emotional dependence and problematic usage compared to open-ended conversations. However, with longer usage duration this effect becomes non-significant. 4. Initial States Pre-existing measures of emotional well-being were statistically significant predictors of post-interaction states. Participants who started with high initial emotional dependence and problematic use had statistically significantly reduction in both measures using the engaging voice modality compared to the text modality.

Limitations. One limitation of surveys is that the results are self-reported by users, and may reflect their selfperception more than their actual behavior or revealed preferences. To compare users’ self-reported responses with their actual usage patterns, we pair our survey analysis with methods for analyzing of user conversation that preserve their privacy. To study the emotional content in user conversations in an automated manner, we run the EmoClassifiersV1 (Section 2) on the conversations of both cohorts within the study period. This While live platform usage provides a rich set of data for analysis, there are significant limitations in the kinds of research questions that can be answered (see also Table 2):

• User Information: The ChatGPT platform currently does not collect a lot of key information about its users that we may like to control for in our analysis, such as gender or prior familiarity with AI. • User Feedback: Beyond usage data, we would also like to get quantitative or qualitative feedback on their experience using models. However, it can be difficult to get users to fill in surveys or provide detailed feedback, and results from voluntarily filled out surveys will be subject to issues of selection bias. • Experimental Constraints: We are unable to dictate usage of a certain model configuration (e.g.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What design features sustain romantic bonds with AI companion systems? How can emotionally responsive AI maintain reliability and healthy boundaries? Can AI chatbots provide mental health support without reinforcing harmful beliefs? What enables conversational agents to guide rather than just respond? Why do confident AI outputs mislead human trust calibration? When do multi-agent systems improve over single frontier models? How does AI-generated content create social proof without authentic interaction? How can agents discover and adapt to user preferences during conversation? How does personalization simultaneously affect user trust and privacy concerns? Do individually safe AI actions create unsafe outcomes in integrated systems? Why do people trust AI chatbots with sensitive information?