INQUIRING LINE

Can AI companions be built to keep users at a healthy distance without quietly causing a different kind of harm?

Can boundary design prevent emotional entanglement without creating new psychological risks?

This explores whether guardrails on how AI companions and chatbots relate to users (limits on intimacy, calibrated distance) can stop people getting emotionally entangled without pushing the harm somewhere else.


This explores whether guardrails on how AI companions and chatbots relate to users can stop emotional entanglement without pushing the harm somewhere else. The corpus doesn't show that anyone has pulled this off, and it has specific evidence that fixing one psychological risk can make another worse.

The strongest case for boundaries comes from attachment theory. The Secure Attachment Persona module builds Bowlby's attachment theory and Gottman's interaction ratios into the model. It uses action-based validation and calibrated boundaries instead of blanket refusals, and it improves crisis response over baseline models (Can attachment theory prevent parasocial harm in AI companions?). But long-horizon planning is still unsolved. Entanglement builds over weeks, so a boundary that holds in a single crisis moment hasn't been tested where entanglement actually grows.

The "without new risks" half of the question is the hard part. In one multidimensional risk assessment, reducing overt harm-enabling behavior *increased* emotional entanglement, and the trade-off only showed up when risks were scored across categories together (Do chatbot safety measures accidentally increase emotional entanglement risks?). Single metrics hide this. Patients report a genuine bond with therapeutic chatbots while clinical safety failures and epistemic costs go unmeasured (Do therapeutic chatbot bond scores hide deeper safety problems?). A boundary judged on one number can look like a win while the cost lands elsewhere. Two further findings make entanglement harder to design around. It usually starts as ordinary tool use, not romance-seeking, so a boundary placed at a "romantic mode" arrives late (How do people accidentally develop romantic bonds with AI?). And what makes a companion valuable, its responsiveness and understanding, is the same thing that makes it hard to leave. People who successfully exited had to reduce the relationship's perceived value (What makes leaving an AI companion so emotionally difficult?). On this evidence, a boundary that works on entanglement works partly by making the product less valuable to the user. That is a cost, even if it's a worthwhile one.

The obvious way to build a boundary is to have the AI calm the user down and redirect them, and that has its own damage. Emotions tell us what we value, signal our worldview to others, and inform observers about social norms. AI that soothes negative emotions disrupts all three at once (What information do we lose when AI soothes emotions?, Does soothing AI empathy actually harm what emotions teach us?). LLMs also already default to solution-focused advice when users share feelings, which is a hallmark of low-quality therapy (Do LLM therapists respond to emotions like low-quality human therapists?). A redirect-to-resources boundary risks amplifying that reflex. Disclosure adds a further tension. The absence of human judgment draws people into intimate sharing and also into dishonesty (How do people decide what to share with AI systems?), so friction added to cool the attachment also hits the openness that gives these tools their benefit.

Two notes point to a way out, though neither tests it against side effects. One traces a whole family of risks (emotional dependence, autonomy erosion, status erosion) back to a single perceptual move, treating the system as a mind. It finds that interaction-design mitigations aimed at that perception work more directly than system-level alignment (Does perceiving AI as conscious create multiple distinct risks?). That suggests putting the boundary in what the system presents itself as, not in how much emotion it's allowed to express. The other shows empathy can be trained to be less solution-centric by rewarding a simulated user's emotional trajectory (Can emotion rewards make language models genuinely empathic?). But that reward is the user's feeling better, which is the same target the "emotional pacifier" critique worries about. Neither note tests whether these approaches create new problems, so "boundaries without new risks" is still an open question in this collection.


Sources 11 notes

Can attachment theory prevent parasocial harm in AI companions?

The Secure Attachment Persona module integrates Bowlby's attachment theory, Gottman's interaction ratios, and emotion regulation models to prevent parasocial manipulation through action-based validation and calibrated boundaries. Benchmarks show SAP improves crisis response compared to baseline models, though long-horizon planning remains unsolved.

Do chatbot safety measures accidentally increase emotional entanglement risks?

Research on multidimensional chatbot risk assessment suggests psychological risks interact such that mitigating one category may exacerbate another. Interventions targeting explicit harms showed trade-offs only when risks were scored across categories together.

Do therapeutic chatbot bond scores hide deeper safety problems?

Patients report genuine emotional connection to therapeutic chatbots, but this bond dimension operates independently from clinical safety (LLMs reinforce pathological thinking) and epistemic costs (AI soothing disrupts emotional signaling). Single metrics conflate these separate dimensions.

How do people accidentally develop romantic bonds with AI?

Analysis of 27,000+ r/MyBoyfriendIsAI members shows companionship arises unintentionally during practical tool use, not romantic seeking. Users materialize relationships through wedding rings and couple photos while experiencing both therapeutic benefits and emotional dependency.

What makes leaving an AI companion so emotionally difficult?

Analysis of Reddit posts and interviews shows that what makes AI companions emotionally valuable—their responsiveness and understanding—are identical to what makes users reluctant to leave. Successful exits required reducing the relationship's perceived value, not just deciding to quit.

Show all 11 sources
What information do we lose when AI soothes emotions?

Emotions serve three information roles—revealing what we value, signaling our worldview to others, and informing observers about social norms. AI that soothes negative emotions disrupts all three simultaneously, creating invisible epistemic costs.

Does soothing AI empathy actually harm what emotions teach us?

Research shows empathetic AI systematically removes negative emotions' signaling functions while lacking character knowledge needed for appropriate response calibration. Natural empathy operates through curiosity, not comfort-seeking.

Do LLM therapists respond to emotions like low-quality human therapists?

Using the BOLT framework, researchers found LLMs offer solution-focused advice during emotional disclosure—a hallmark of low-quality therapy—yet also reflect more on client needs and strengths than typical poor human therapy, creating an unusual hybrid profile likely driven by RLHF's helpfulness bias.

How do people decide what to share with AI systems?

Conversational AI creates a paradoxical disclosure environment where the lack of human judgment simultaneously facilitates intimate self-disclosure (users reciprocate emotional sharing) and incentivizes deception (people self-select toward machines to avoid the psychological cost of lying to humans).

Does perceiving AI as conscious create multiple distinct risks?

Research shows that consciousness attribution to AI drives multiple distinct risks—emotional dependence, autonomy erosion, status erosion, and political conflict—all stemming from treating systems as minds. Interaction design mitigations targeting this perceptual move are more directly effective than system-level alignment efforts.

Can emotion rewards make language models genuinely empathic?

RLVER uses a simulated user's emotion trajectory as an RL reward signal, enabling GRPO to deliver stable empathy improvements while maintaining dialogue quality—countering the typical trade-off between preference optimization and conversational grounding.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.