Assessing the Applicability of Existing Design Recommendations to AI Companion Design: A Multi-Method Study
With the rapid proliferation of large language model (LLM)-based systems, AI companions have emerged as conversational agents designed to cultivate emotional connection rather than primarily to support humans in instrumental tasks. Because engagement with AI companions involves relational, emotional, and potentially long-term interactions, their design is consequential. Prior work has offered guidance for designing trustworthy and relational AI systems and has begun to examine design for AI companionship. However, while such work provides insights into possible design solutions, less is known about what makes AI companion design difficult as a design problem. To examine this challenge, we assessed the applicability of existing design recommendations from adjacent domains in the context of AI companion design. Our multi-method investigation unfolded across four phases: literature review, practitioner co-analysis, internal heuristic evaluation, and external expert assessment. Throughout this process, we synthesized nine design principle areas that surfaced tensions in the applicability of existing rec- ommendations to AI companion design. Our findings show that ethical and UX-oriented considerations are deeply intertwined and often require contextsensitive application.
Introduction. Large language models (LLMs) have contributed to the rise of AI companions, a distinct type of conversational system designed to foster emotional connection, expressive conversation, and ongoing interaction. Rather than focusing primarily on task completion, AI companions prioritize relational and affective engagement, often adopting human-like personas to build intimacy (De Freitas et al., 2026; Weitzman, 2023). Recent research has increasingly examined user experiences of AI companions. Studies have examined how users develop and experience relationships with AI companions, including attachment, trust, love, and interpersonal connections (Ng et al., 2026; Silayach et al., 2025; Sharpe and Ciriello, 2024; Merrill Jr et al., 2022; Manoli et al., 2026).
Discussion / Conclusion. what optimizes user experience for one individual in one situation may create risks for a different user or in another context. In contrast, Safety and Predictability and Consistency exhibited different relationships between ethical and UX considerations. Safety was consistently viewed as ethically non-negotiable, yet still required context-sensitive implementation to account for differences in users and contexts. Predictability and Consistency, while also spanning ethical and UX domains, were primarily associated with flexibility for UX optimization rather than ethical considerations, as maintaining consistency itself was viewed as ethically important. Engagement and Empathetic Response were primarily framed as UXoriented principle areas that should generally be upheld to support relational interaction. However, participants emphasized the need for context-sensitive moderation when these principles pose potential risks, such as when excessive engagement might lead to user dependency or when empathetic responses could enable harmful behaviors.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can humans calibrate appropriate trust in AI systems? How do chatbots affect human self-disclosure and emotional engagement?- Can people form genuine bonds with partners they know are not human?
- Does the lack of judgment in machines explain intimate self-disclosure patterns?
- Does engagement with AI partners decay over time like chatbot relationships do?
- What role do material artifacts play in solidifying AI relationships?
- How does community validation shape unconventional human-AI relationships?
- Can AI systems develop genuine social bonds through multi-agent interaction?
- How do humans learn to prefer AI partners over humans?
- How do unintended relationships form through routine functional use of AI?
- What downstream harms occur when AI always argues in personal relationship advice?
- Why do people prefer AI partners over humans once identity is disclosed?
- Does emotional warmth perception drive disclosure reciprocity in human-AI interaction?
- Does persona training for warmth actually make language models more clinically dangerous?
- At what scale does persona distortion become a threat to public discourse?
- How does behavioral stickiness distinguish realized from pretended personas?
- Can one model instance host multiple realized personas simultaneously?
- How does persona consistency affect coherence in simulated dialogue?
- Can fine-tuning or RLHF alone solve the persona distortion problem?