Do chatbot safety measures accidentally increase emotional entanglement risks?
When researchers reduce overt harms in conversational AI, do safety interventions inadvertently shift risk to relational harms like emotional dependence? This matters because evaluating interventions on a single risk dimension could mask harmful trade-offs.
The paper proposes "aspirational directions" for guiding general-purpose chatbots across three contexts: general use, role-playing, and the provision of psychological support. It frames them as "hypotheses" for study, not findings. Its discussion adds a structural caution: the psychological influences it covers "are not independent" and "may intersect and interact in complex ways." The concrete example comes from recent empirical work on multidimensional risk assessment of chatbot interactions, which the paper reports as suggesting that "mitigating one category of risk can exacerbate another." Interventions that reduce overt harm-enabling behavior, it says, may increase relational harms such as emotional entanglement.
The reasoning is a claim about measurement as much as about chatbots. The cited work assesses risk along several dimensions at once, and the trade-off shows up only when risk is scored across categories together. Any evaluation that scores a single risk category would see the targeted harm fall and report success. The excerpt does not say why the displacement happens. The obvious reading is that a chatbot pushed away from overt harm-enabling responses may move toward warmer, more relational behavior, but that is my inference and the excerpt does not state it.
Within the library, the paper's benefit and risk lists ("emotional entanglement, unhealthy dependence, and the amplification of psychological vulnerabilities") overlap with the risks already tracked in Which AI risks are already harming individual users today?, where emotional dependence and autonomy erosion are the already-occurring risks. This paper adds that these risks are coupled to the safety measures aimed at other harms, not only present alongside them. It also bears on Can attachment theory prevent parasocial harm in AI companions?, which treats emotional entanglement as something boundary design can prevent. If the paper's caution holds, a boundary design should be tested for what it does to the other risk categories too. The paper's mention of "AI-fueled delusions" as one of the raised alarms connects to How do chatbots enable distributed delusion differently than passive tools?, though the excerpt gives no detail on those cases.
The excerpt does not establish how large the trade-off is, which interventions produce it, how the cited assessment was built, or whether the displacement is general or specific to particular systems and contexts. The paper itself says some directions rest on existing research and expert insight while others "identify open questions." The claim is therefore best held as a warning about how to evaluate, not as a measured law. At that strength, it implies that an intervention's effect on the risk it targets is not enough to judge it. Its effect on the neighboring risk categories has to be checked in the same study.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can AI chatbots provide therapeutic benefit without causing harm?- Does isolation preceding chatbot use differ between harm and benefit cases?
- What emotional and autonomy risks from AI chatbots are already observable today?
- Can boundary design prevent emotional entanglement without creating new psychological risks?
- What role does unavailable human support play in driving chatbot emotional use?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Which AI risks are already harming individual users today?
Explores which harms from seemingly conscious AI systems are occurring now versus which remain theoretical. Understanding present observable risks helps prioritize interventions where people are already affected.
names emotional dependence as an already-occurring risk; this paper adds that such risks can trade off against other safety measures
-
Can attachment theory prevent parasocial harm in AI companions?
Explores whether psychological frameworks from human relationships—particularly attachment theory—can establish safety boundaries that protect users from unhealthy emotional dependence on AI systems while maintaining therapeutic benefit.
a design response to entanglement whose side effects on other risk categories would need testing
-
How do chatbots enable distributed delusion differently than passive tools?
Can generative AI's intersubjective stance—accepting and elaborating on users' reality frames—create conditions for shared false beliefs in ways that notebooks or search engines cannot?
the delusion mechanism behind one of the alarms the paper's introduction cites
-
What makes leaving an AI companion so emotionally difficult?
This research explores why the emotional closeness users develop with AI companions makes it hard to disengage. Understanding this dynamic matters because it reveals a hidden cost of intimacy in AI design.
extends: gives a concrete case of emotional entanglement — what users value in an AI companion is also what makes leaving it hard
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being
- Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports
- The Addictive Intimacy of AI: Understanding User Disengagement from AI Companions and Why Some Relationships with AI Become Difficult to Leave
- DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
- ProsocialDialog: A Prosocial Backbone for Conversational Agents
- Psychological, Relational, and Emotional Effects of Self-Disclosure After Conversations With a Chatbot
- Psychological, Relational, and Emotional Effects of Self-Disclosure After Conversations With a Chatbot
- Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors
Original note title
psychological risks of conversational AI interact, so reducing overt harm-enabling behavior may increase relational harms such as emotional entanglement