Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors
Ruptures represent common albeit critical moments in interaction where relational alignment breaks down, making them essential for evaluating AI where trust and engagement matter most. In a scenario-driven empirical study, we examined the performance of three LLMs at identifying and resolving ruptures across 21 mental health conversations and 22 experts’ evaluation of the strategies. For identification, LLMs relied on explicit linguistic cues within single turns whereas experts integrated implicit, relational, and contextual information across the conversation. For resolution, LLMs tended to produce more directive and scripted responses whereas experts adopted process-oriented strategies such as validation, open-ended exploration, and psychoeducation. Overall, LLMs showed higher agreement with predefined labels in identification, but not in resolution where experts rated their responses only moderately effective, with consistent limitations in timing, depth, and contextual sensitivity. We discuss implications for the design of mental health conversational agents emphasizing relational awareness, pacing, and human-in-the-loop support.
Introduction. Large language model-enabled AI conversational systems (chatbots) are increasingly used in emotionally sensitive and high-stakes contexts, including mental health support [16, 45, 75, 79], caregiver guidance [31, 67, 68], patient-facing healthcare [18, 78], and personal reflection, such as journaling and self-tracking [17, 42, 43]. In these settings, users turn to AI systems for more than information, including emotional support, guidance, and reassurance [29, 38, 69, 76]. As a result, interaction effectiveness depends on the quality of the perceived relationship between users and AI systems. Prior work identifies this relationship as central to user engagement and trust [8, 20, 69, 76]. Development has moved well beyond clinical research teams. General-purpose assistants, lightweight application wrappers, and configurable personas have lowered the barriers to deploying a system that functions as mental health support, often without clinical training, institutional oversight, or an evaluation plan [16, 34, 79].
Discussion / Conclusion. High agreement on rupture identification should not be interpreted as evidence of clinical understanding. As Sections 5.1, 5.2, and 5.3 show, LLMs often matched reference rupture labels but relied heavily on explicit linguistic cues rather than deeper interpretation of relational dynamics. This gap was clearer in resolution and evaluation tasks. Experts rated the models’ resolution strategies as only moderately effective and identified recurring problems, including premature problem solving, shallow engagement, and mechanical tone. Clinical interpretation therefore requires more than selecting the reference category. Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors 21 Table 7. experts’ broader perceptions of AI chatbots in mental health counseling mapped to rupture-related concerns and resolution implications.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do chatbots affect human self-disclosure and emotional engagement?- How does consciousness attribution drive emotional dependence on chatbots?
- How does emotional dependence on chatbots affect user wellbeing?
- Can people form genuine bonds with partners they know are not human?
- What harms might chatbots cause through stigma expression and delusion reinforcement?
- Why do therapists and patients report misaligned perceptions of the working relationship?
- Can real-time therapist feedback improve outcomes using computational alliance measurement?
- Can single-turn empathy advantage predict multi-turn therapeutic outcomes?
- What separates generating empathic responses from maintaining therapeutic alliance?
- How does turn-level working alliance inference enable real-time therapist feedback?
- Does true understanding matter for therapeutic benefits of disclosure?
- Can people form therapeutic bonds with tools they know are not human?
- What clinical harms might hide behind positive therapeutic bond measurements?
- Can therapeutic bonds exist without genuine reciprocity or mutual understanding?
- How do language models interpolate user feelings in therapeutic contexts?
- How should AI systems separate feeling interpretation from objective therapeutic guidance?
- Why do mental health chatbots fail at synchrony despite strong language models?
- Do therapeutic chatbots adequately detect crisis situations and safety risks?
- How do dropout rates and low adherence affect chatbot therapy outcomes?
- What architectural changes would enable proactive therapeutic guidance in chatbots?