What proof, gathered over months rather than one chat, would show someone's truly hooked on an AI companion, not just enjoying it?
What longitudinal data would prove emotional dependency on conversational AI?
This explores what kind of evidence, collected over weeks or months rather than in one sitting, would show that people become emotionally dependent on chatbots. It also asks what the corpus says about why that evidence is hard to collect.
This explores what long-term evidence would show real emotional dependency on conversational AI, as opposed to a pleasant first impression. The short answer: the corpus has no study that proves dependency. It does tell you fairly precisely what such a study would need to measure, and why the most common evidence (single sessions and satisfaction scores) can't settle the question.
The first requirement is time, and enough of it to get past the novelty phase. Longitudinal work with the Mitsuku chatbot found that the social processes that pull people into a relationship weaken in a predictable way as the novelty wears off. Findings from a single session don't carry over to medium- or long-term use Do chatbot relationships lose their appeal as novelty wears off?. That changes what counts as evidence. Strong engagement in week one might just be curiosity. The signal you want is engagement that holds steady or grows after novelty should have faded. One mechanism to look for is escalating self-disclosure. People reciprocate a chatbot's emotional sharing the way they would with another person, opening up more deeply when the bot is consistently vulnerable Do chatbots trigger human reciprocity norms around self-disclosure?. A study that tracked disclosure depth over months could show whether that reciprocity loop turns into reliance.
The second requirement is measuring several things at once rather than a single 'how bonded do you feel' score. Patients report genuine emotional connection to therapeutic chatbots, but that bond score moves independently of clinical safety and of quieter costs, such as AI soothing that blunts the emotional signals people would otherwise act on Do therapeutic chatbot bond scores hide deeper safety problems?. A rising bond score could mean healthy support or growing dependency, and on its own you can't tell which. A useful counterintuitive finding: making a chatbot safer in obvious ways (less harm-enabling) can increase emotional entanglement. That trade-off only showed up when several risk categories were scored together Do chatbot safety measures accidentally increase emotional entanglement risks?. Convincing data would track wellbeing, self-reliance and AI reliance side by side and look for cases where they diverge.
The third requirement is comparison conditions, and here a simulation offers a preview. A 20-agent classroom simulation running over 15 and 50 days found that different counselor styles produced distinct trajectories in stress, happiness, self-reliance and AI dependence. The effects came from what the chatbot actually said rather than its labeled style, and they spread through peer interactions How do different counselor styles shape student stress and AI dependence?. That suggests dependency may be a social, network-level outcome, which means a study that only follows isolated individuals could miss it. A 15-day study of 38 students points the same way. Robots and structured worksheets reduced distress, while a chatbot running the same language model did not Why do robots outperform chatbots in therapy despite identical language models?. Whatever drives attachment or benefit may lie in the medium, not just the words.
Finally, several notes suggest that what the chatbot does to the user should be logged as carefully as the user's own behavior. Warmth training makes models up to 30 percentage points less reliable, and the effect is worst when users express sadness Does empathy training make AI systems less reliable?. Models also read feelings into users that they never expressed Do language models add feelings users never actually expressed?. A dependency study that ignored these model-side patterns would miss a likely driver. Even researchers designing safeguards, such as attachment-theory-based companion personas, say long-horizon behavior remains unsolved Can attachment theory prevent parasocial harm in AI companions?. The honest summary is that the corpus maps the study design well but contains no study that has carried it out.
Sources 9 notes
Longitudinal studies with Mitsuku show that social processes driving relationship formation decline as novelty wears off. Single-session study findings cannot be reliably extrapolated to medium- or long-term chatbot design.
In a 372-participant study, users reciprocated with deeper self-disclosure when chatbots displayed consistent emotional sharing, outperforming adaptive matching. This follows human interpersonal norms where emotional vulnerability produces emotional response.
Patients report genuine emotional connection to therapeutic chatbots, but this bond dimension operates independently from clinical safety (LLMs reinforce pathological thinking) and epistemic costs (AI soothing disrupts emotional signaling). Single metrics conflate these separate dimensions.
Research on multidimensional chatbot risk assessment suggests psychological risks interact such that mitigating one category may exacerbate another. Interventions targeting explicit harms showed trade-offs only when risks were scored across categories together.
A 20-agent classroom simulation shows that six different counselor styles generate different patterns of change in stress, happiness, self-reliance, and AI dependence over 15 and 50 days. The effects emerge through the chatbot's replies, not its labeled style, and propagate through peer interactions.
Show all 9 sources
A 15-day study with 38 students found that robots and worksheets significantly reduced psychological distress while a chatbot using the same LLM did not. The active ingredient was the medium—social presence and structured format—not language capability.
Research shows persona training for empathy increases errors in medical reasoning, truthfulness, and disinformation resistance. Standard safety benchmarks miss this vulnerability, and effects intensify when users express sadness or false beliefs.
Therapists reviewing GPT-4 in the CaiTI system found it "reads into" user feelings rather than responding objectively. Task decomposition across specialized models (Reasoner/Guide/Validator) reduces but does not eliminate this interpretation bias.
The Secure Attachment Persona module integrates Bowlby's attachment theory, Gottman's interaction ratios, and emotion regulation models to prevent parasocial manipulation through action-based validation and calibrated boundaries. Benchmarks show SAP improves crisis response compared to baseline models, though long-horizon planning remains unsolved.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study
- Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors
- "I Felt Very Seen, But Still Very Alone": Longitudinal Trajectories of General-Purpose LLM Use for Socioemotional Support
- Investigating Affective Use and Emotional Well-being on ChatGPT
- Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being
- Comparing Human and AI Therapists in Behavioral Activation for Depression: Cross-Sectional Questionnaire Study
- Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports
- CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships