"I Felt Very Seen, But Still Very Alone": Longitudinal Trajectories of General-Purpose LLM Use for Socioemotional Support
People increasingly use general-purpose chatbots such as ChatGPT, Claude, and Gemini for mental health and emotional support. We report a multi-stage longitudinal qualitative study of 18 U.S. adults, conducted from April to December 2025, combining initial interviews, a four-week diary study, focus groups, and exit interviews. We find that socioemotional use often emerged gradually out of practical use and when other forms of support were unavailable. Participants developed routines and boundaries around chatbot use, which were disrupted by model updates, evolving public discourse about AI harms, and changes in personal circumstances. We demonstrate how longitudinal study captures factors beyond the human-AI dyad, and argue that HCI researchers and designers should account for users’ histories with their chatbots and broader care ecologies when evaluating AI systems over time and introducing updates that may disrupt established sources of support.
Introduction. General-purpose large language model (LLM) chatbots such as ChatGPT, Claude, and Gemini have become a source of mental health and emotional support despite not being initially designed for that purpose and operating largely outside clinical regulatory frameworks [2, 27, 45]. Nearly one in five U.S. adolescents and young adults reported using AI chatbots for mental health advice [47], while among adults with a diagnosed mental health condition who use LLMs, close to half reported turning to them for therapeutic or emotional support [65]. Platform-scale analyses likewise identify affective use across varying levels of engagement, although it is concentrated among a minority of heavy users [63]. We study how a general-purpose chatbot comes to occupy a socioemotional role in a person’s life, and what changes that role over time. Several conditions help explain why people turn to AI chatbots for support.
Discussion / Conclusion. This study demonstrates that responsibly contextualizing socioemotional interactions with chatbots requires examining the human-chatbot relationship beyond a single conversation and over time. Below, we highlight key aspects and recommendations for the HCI community to consider when studying general-purpose chatbots for affective use. 5.1 Grounding Chatbot Evaluation in Lived Experience Grounding evaluations in user experiences helps to identify effects of chatbot use that are not captured through current methods. Existing evaluation frameworks in this space are built primarily from technical categories and rarely include the input of those who use AI chatbots for mental and emotional support [4, 20, 78]. This approach prioritizes a top-down approach, therefore limiting how accurately evaluations can capture what users experience as supportive or harmful. In our own research, P13’s account of feeling seen while remaining alone illustrates why perceived understanding and a sense of connection should be separately examined (§4.3). Examples such as these illustrate how users’ accounts can help specify how metrics should be categorized, defined, and prioritized [15, 53].
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can LLM user simulators model realistic goal-driven conversation? Why do LLM chatbots fail as independent therapeutic agents?- Can trainees improve formulation skills by practicing against simulated patients?
- How do language models interpolate user feelings in therapeutic contexts?
- Why do mental health chatbots fail at synchrony despite strong language models?
- Do therapeutic chatbots adequately detect crisis situations and safety risks?
- How do dropout rates and low adherence affect chatbot therapy outcomes?
- What architectural changes would enable proactive therapeutic guidance in chatbots?
- How do waitlist-control RCTs mislead about therapeutic chatbot real-world efficacy?
- Does social presence from robots drive adherence better than conversational AI interfaces?
- Do worksheet-based structured formats work as well as embodied agents for therapy?
- Do conversational AI systems overuse first-person pronouns in therapy settings?
- What role does conversational presence play in making therapy feel reciprocal?
- How does RLHF training push therapeutic chatbots toward problem-solving over attunement?
- Can architectural constraints on model input reduce emotional interpolation in clinical AI?
- What separates generating empathic responses from maintaining therapeutic alliance?
- Can people form therapeutic bonds with tools they know are not human?
- How do bond scores predict actual therapy outcomes in digital interventions?