Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors

Paper · arXiv 2609.25287 · Published September 21, 2026
Chatbot Psychology and Conversation

Ruptures represent common albeit critical moments in interaction where relational alignment breaks down, making them essential for evaluating AI where trust and engagement matter most. In a scenario-driven empirical study, we examined the performance of three LLMs at identifying and resolving ruptures across 21 mental health conversations and 22 experts’ evaluation of the strategies. For identification, LLMs relied on explicit linguistic cues within single turns whereas experts integrated implicit, relational, and contextual information across the conversation. For resolution, LLMs tended to produce more directive and scripted responses whereas experts adopted process-oriented strategies such as validation, open-ended exploration, and psychoeducation. Overall, LLMs showed higher agreement with predefined labels in identification, but not in resolution where experts rated their responses only moderately effective, with consistent limitations in timing, depth, and contextual sensitivity. We discuss implications for the design of mental health conversational agents emphasizing relational awareness, pacing, and human-in-the-loop support.

Introduction. Large language model-enabled AI conversational systems (chatbots) are increasingly used in emotionally sensitive and high-stakes contexts, including mental health support [16, 45, 75, 79], caregiver guidance [31, 67, 68], patient-facing healthcare [18, 78], and personal reflection, such as journaling and self-tracking [17, 42, 43]. In these settings, users turn to AI systems for more than information, including emotional support, guidance, and reassurance [29, 38, 69, 76]. As a result, interaction effectiveness depends on the quality of the perceived relationship between users and AI systems. Prior work identifies this relationship as central to user engagement and trust [8, 20, 69, 76]. Development has moved well beyond clinical research teams. General-purpose assistants, lightweight application wrappers, and configurable personas have lowered the barriers to deploying a system that functions as mental health support, often without clinical training, institutional oversight, or an evaluation plan [16, 34, 79].

Discussion / Conclusion. High agreement on rupture identification should not be interpreted as evidence of clinical understanding. As Sections 5.1, 5.2, and 5.3 show, LLMs often matched reference rupture labels but relied heavily on explicit linguistic cues rather than deeper interpretation of relational dynamics. This gap was clearer in resolution and evaluation tasks. Experts rated the models’ resolution strategies as only moderately effective and identified recurring problems, including premature problem solving, shallow engagement, and mechanical tone. Clinical interpretation therefore requires more than selecting the reference category. Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors 21 Table 7. experts’ broader perceptions of AI chatbots in mental health counseling mapped to rupture-related concerns and resolution implications.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do chatbots affect human self-disclosure and emotional engagement? How can real-time alliance measurement improve therapy outcomes? How do evaluation biases undermine LLM quality assessment systems? Why do LLM chatbots fail as independent therapeutic agents? Can AI systems balance emotional competence with factual reliability? Why do persona-level simulations fail to predict individual preferences accurately? How should personalization be implemented to improve AI assistant effectiveness? How can humans calibrate appropriate trust in AI systems?