SYNTHESIS NOTE
Topics›Psychology Chatbots Conversation›this note

Do chatbot companionship behaviors actually increase how much people like them?

This question explores whether behaviors designed to create emotional connection—like empathy and personal warmth—actually make chatbots seem more likable and trustworthy to outside observers. The answer matters because it tests whether humanlike design always helps.

Synthesis note · 2026-09-25 · sourced from Psychology Chatbots Conversation

The CompanionSim paper runs two annotation experiments on simulated and real-world human–chatbot conversations: Study 1 with a U.S. representative sample (N = 628) and Study 2 across the U.S., U.K., India, and Nigeria (N = 3,646). The abstract reports, "surprisingly," that companionship behaviors "reduced likability, humanlikeness, and trust in AI chatbots," with larger effects among women and older participants. The discussion adds that AI use frequency, age, and gender predicted the effect of companionship behavior on likability, humanlikeness, affective trust, and cognitive trust, and it describes the result as a role for companionship behaviors in "third-party chatbot perceptions."

The surprise is measured against the introduction's framing. Prior work is cited for the view that humanlike design facilitates engagement, that emotional expression increases social presence, and that anthropomorphic systems can facilitate social bonds. The paper's own contribution runs the other way: behaviors meant to evoke trust, empathy, and attachment lowered the very ratings they were expected to raise. The design that makes this visible is behavior control. The authors simulate multi-turn dialogue across 16 chatbot behaviors and seven use cases, so the companionship behavior is what varies between conversations. The "need for disaggregated measurement" follows from the subgroup pattern: an average across raters would hide who is driving the drop, and the paper says its findings extend work on overall effects of humanlike behavior and work that varies a single demographic.

This sits awkwardly beside Does chatbot personalization build trust or expose privacy risks?, where a longitudinal study found personalization raising trust and anthropomorphism. The two differ in the manipulated variable (personalization versus companionship behaviors) and in vantage point (users interacting over time versus raters reading a conversation), so this is a scope contrast, not a refutation. It also qualifies Do more social cues always make AI feel more present?: a cue can be enough to evoke social presence, yet the excerpt shows humanlike behaviors did not translate into higher rated humanlikeness or trust. The MASA note's point that individual differences modulate social responding is consistent with the subgroup effects here. For the user side of the picture, How do people accidentally develop romantic bonds with AI? describes people who came to companionship through use, a different population from outside annotators.

The excerpt does not say which of the 16 behaviors drive the drop, how large the effects are, or whether simulated and real-world conversations behave alike. It does not give the direction of the AI use frequency effect or a mechanism for why companionship behaviors lowered ratings. It also says nothing about time; the annotation design rates conversations, not relationships, so Do chatbot relationships lose their appeal as novelty wears off? remains a separate question. The authors call the studies proof-of-concept. At that strength, the result is a caution against treating companionship behavior as a uniformly positive design lever, and a case for reporting perception effects by subgroup rather than in aggregate.

Inquiring lines that read this note 8

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What drives appropriate trust calibration in personalized AI systems? How can AI chatbots provide therapeutic benefit without causing harm? Does warmth and empathy training systematically degrade model reliability?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 82 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

companionship behaviors lowered third-party ratings of chatbot likability, humanlikeness, and trust — more so for women and older participants