SYNTHESIS NOTE
Topics›Psychology Chatbots Conversation›this note

Do chatbot safety measures accidentally increase emotional entanglement risks?

When researchers reduce overt harms in conversational AI, do safety interventions inadvertently shift risk to relational harms like emotional dependence? This matters because evaluating interventions on a single risk dimension could mask harmful trade-offs.

Synthesis note · 2026-09-25 · sourced from Psychology Chatbots Conversation

The paper proposes "aspirational directions" for guiding general-purpose chatbots across three contexts: general use, role-playing, and the provision of psychological support. It frames them as "hypotheses" for study, not findings. Its discussion adds a structural caution: the psychological influences it covers "are not independent" and "may intersect and interact in complex ways." The concrete example comes from recent empirical work on multidimensional risk assessment of chatbot interactions, which the paper reports as suggesting that "mitigating one category of risk can exacerbate another." Interventions that reduce overt harm-enabling behavior, it says, may increase relational harms such as emotional entanglement.

The reasoning is a claim about measurement as much as about chatbots. The cited work assesses risk along several dimensions at once, and the trade-off shows up only when risk is scored across categories together. Any evaluation that scores a single risk category would see the targeted harm fall and report success. The excerpt does not say why the displacement happens. The obvious reading is that a chatbot pushed away from overt harm-enabling responses may move toward warmer, more relational behavior, but that is my inference and the excerpt does not state it.

Within the library, the paper's benefit and risk lists ("emotional entanglement, unhealthy dependence, and the amplification of psychological vulnerabilities") overlap with the risks already tracked in Which AI risks are already harming individual users today?, where emotional dependence and autonomy erosion are the already-occurring risks. This paper adds that these risks are coupled to the safety measures aimed at other harms, not only present alongside them. It also bears on Can attachment theory prevent parasocial harm in AI companions?, which treats emotional entanglement as something boundary design can prevent. If the paper's caution holds, a boundary design should be tested for what it does to the other risk categories too. The paper's mention of "AI-fueled delusions" as one of the raised alarms connects to How do chatbots enable distributed delusion differently than passive tools?, though the excerpt gives no detail on those cases.

The excerpt does not establish how large the trade-off is, which interventions produce it, how the cited assessment was built, or whether the displacement is general or specific to particular systems and contexts. The paper itself says some directions rest on existing research and expert insight while others "identify open questions." The claim is therefore best held as a warning about how to evaluate, not as a measured law. At that strength, it implies that an intervention's effect on the risk it targets is not enough to judge it. Its effect on the neighboring risk categories has to be checked in the same study.

Inquiring lines that read this note 7

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can AI chatbots provide therapeutic benefit without causing harm? Does warmth and empathy training systematically degrade model reliability? Why do locally safe actions create system-level safety gaps?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 77 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

psychological risks of conversational AI interact, so reducing overt harm-enabling behavior may increase relational harms such as emotional entanglement