Large language models could change the future of behavioral healthcare: a proposal for responsible development and evaluation
LLMs and psychotherapy skills For certain use cases, LLM show a promising ability to conduct tasks or skills needed for psychotherapy, such as conducting assessment, providing psychoeducation, or demonstrating interventions (see Fig. 2). Yet to date, clinical LLM products and prototypes have not demonstrated anywhere near the level of sophistication required to take the place of psychotherapy. For example, while an LLM can generate an alternative belief in the style of CBT, it remains to be seen whether it can engage in the type of turn-based, Socratic questioning that would be expected to produce cognitive change. This more generally highlights the gap that likely exists between simulating therapy skills and implementing them effectively to alleviate patient suffering. Given that psychotherapy transcripts are likely poorly represented in the training data for LLMs, and that privacy and ethical concerns make such representation challenging, prompt engineering may ultimately be the most appropriate fine-tuning approach for shaping LLM behavior in this manner.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do LLM chatbots fail as independent therapeutic agents?- Can models succeed at mental health tasks without integrating multiple psychological traditions?
- What makes Beck's diagram effective for constraining simulated patient behavior?
- Can trainees improve formulation skills by practicing against simulated patients?
- Why do Llama-based models outperform GPT-4 in objective clinical guidance?
- Do problem-solving defaults in LLM therapists actually undermine therapeutic effectiveness?
- Can simulated therapy practice transfer to real-world interpersonal situations?
- What makes clinical theory grounding more effective than pattern matching alone?
- Why do Llama models struggle with cognitively distorted user expressions in therapy?
- Why do LLMs understand therapy techniques but fail to execute them?
- Why can't language models conduct genuine Socratic questioning in therapy sessions?
- Can language models implement therapeutic skills like Socratic questioning in real conversations?
- How does linguistic synchrony differ between LLMs and human therapists over time?
- How do language models interpolate user feelings in therapeutic contexts?
- How do structured cognitive models prevent repetitive and contradictory patient dialogue?
- Why does content richness matter more than linguistic style in patient simulation?