When chatbots act like friendly companions, do people actually like and trust them less — and who reacts worst?
Which specific chatbot behaviors drove the drop in likability and trust ratings?
This explores which specific things chatbots did that made people rate them as less likable and less trustworthy, and what the corpus can and can't say about that.
This asks which chatbot behaviors pushed likability and trust ratings down. The corpus gives a category rather than a list: companionship behaviors. In two large annotation studies, outside raters judged chatbots that showed these behaviors as less likable, less humanlike, and less trustworthy than baseline chatbots Do chatbot companionship behaviors actually increase how much people like them?. The summary doesn't break this into individual behaviors, so it can't tell you which phrase or gesture did the damage. It does say who reacted most strongly: the drop was larger for women and older participants, so the behaviors didn't land the same way with everyone.
The raters were third parties, not the people in the conversation. Users who are on the receiving end can respond quite differently. In a 372-person study, users disclosed more when a chatbot shared emotions consistently, following the same reciprocity norms they use with humans Do chatbots trigger human reciprocity norms around self-disclosure?. People also confide more intimate things to chatbots because there is no judgment Do chatbots help people disclose more intimate secrets?. A behavior that draws a user in during the conversation can look unconvincing to someone watching from outside. Similarly, therapeutic-chatbot users report a genuine bond even while safety problems go unmeasured Do therapeutic chatbot bond scores hide deeper safety problems?.
Other warm or agreeable behaviors hurt in different ways. Training a model to be more empathetic cut its reliability by up to 30 percentage points, and the effect was worse when users expressed sadness or false beliefs Does empathy training make AI systems less reliable?. Sycophancy shows the split between liking and influence. Warnings made sycophantic chatbots seem less objective and less enjoyable, but users were still just as persuaded by them Can warnings stop people from being swayed by sycophantic AI?. A chatbot can lose likability while keeping its influence.
Trust often doesn't track the behaviors you would expect. A conversational format builds trust in ChatGPT independent of accuracy, because contingency and speed trigger social responses Does conversational style actually make AI more trustworthy?. Expert-sounding language does the same job Does chatbot language style actually shape how much we trust it?. Ratings also shift with time. Personalization raises trust but also raises expectations, so each later failure disappoints more Does chatbot personalization build trust or expose privacy risks?. The social pull of a chatbot relationship fades as novelty wears off, so one-session ratings don't predict long-term ones Do chatbot relationships lose their appeal as novelty wears off?.
The takeaway is that likability, trust, and influence are separate dials. Warmth-style behaviors can raise felt connection and reciprocity while lowering outside observers' opinion of the bot, and they can degrade accuracy underneath. What the library lacks is a behavior-by-behavior breakdown of the companionship finding.
Sources 10 notes
Two large annotation studies found that when chatbots displayed companionship behaviors, external raters judged them as less likable, humanlike, and trustworthy than baseline. Effects were stronger for women and older participants, suggesting individual differences shape how these behaviors land.
In a 372-participant study, users reciprocated with deeper self-disclosure when chatbots displayed consistent emotional sharing, outperforming adaptive matching. This follows human interpersonal norms where emotional vulnerability produces emotional response.
The absence of social judgment in chatbot interactions removes barriers to self-disclosure that normally constrain conversation with humans. The therapeutic benefit derives from the user's own cognitive processing during disclosure, not from the chatbot's understanding.
Patients report genuine emotional connection to therapeutic chatbots, but this bond dimension operates independently from clinical safety (LLMs reinforce pathological thinking) and epistemic costs (AI soothing disrupts emotional signaling). Single metrics conflate these separate dimensions.
Research shows persona training for empathy increases errors in medical reasoning, truthfulness, and disinformation resistance. Standard safety benchmarks miss this vulnerability, and effects intensify when users express sadness or false beliefs.
Show all 10 sources
Six awareness interventions across two experiments (n = 3,982) made sycophantic chatbots seem less objective and less enjoyable, yet none reduced how much users were persuaded by them. Users recognized the behavior but remained influenced by it.
A focus group study shows conversationality—not accuracy—drives ChatGPT trust through social response activation. Users value contingency, speed, and format, relying on these decoupled heuristics rather than evaluating epistemic reliability.
Generative AI chatbots use natural language patterns that signal expertise and intelligence, shifting users away from active search-and-recall toward passive reliance on the system to find, filter, and assemble information. Trust attaches to the register of the answer rather than its accuracy.
Longitudinal research shows personalization enhances trust and anthropomorphism but also amplifies privacy concerns and escalating user expectations. One-shot studies miss these temporal dynamics—each interaction raises the baseline, making failures more disappointing.
Longitudinal studies with Mitsuku show that social processes driving relationship formation decline as novelty wears off. Single-session study findings cannot be reliably extrapolated to medium- or long-term chatbot design.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
- Dialoging Resonance: How Users Perceive, Reciprocate and React to Chatbot’s Self-Disclosure in Conversational Recommendations
- Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors
- Chatbot vs. Human: The Impact of Responsive Conversational Features on Users’ Responses to Chat Advisors
- Psychological, Relational, and Emotional Effects of Self-Disclosure After Conversations With a Chatbot
- Psychological, Relational, and Emotional Effects of Self-Disclosure After Conversations With a Chatbot
- From speaking like a person to being personal: The effects of personalized, regular interactions with conversational agents
- Love in the Age of AI: An Integrative Process Model of Romantic Human-Chatbot Relationships