Large language models could change the future of behavioral healthcare: a proposal for responsible development and evaluation

Paper · Source
Therapy Practice and AISynthetic Dialogue Generation

LLMs and psychotherapy skills For certain use cases, LLM show a promising ability to conduct tasks or skills needed for psychotherapy, such as conducting assessment, providing psychoeducation, or demonstrating interventions (see Fig. 2). Yet to date, clinical LLM products and prototypes have not demonstrated anywhere near the level of sophistication required to take the place of psychotherapy. For example, while an LLM can generate an alternative belief in the style of CBT, it remains to be seen whether it can engage in the type of turn-based, Socratic questioning that would be expected to produce cognitive change. This more generally highlights the gap that likely exists between simulating therapy skills and implementing them effectively to alleviate patient suffering. Given that psychotherapy transcripts are likely poorly represented in the training data for LLMs, and that privacy and ethical concerns make such representation challenging, prompt engineering may ultimately be the most appropriate fine-tuning approach for shaping LLM behavior in this manner.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do LLM chatbots fail as independent therapeutic agents? How can LLM user simulators model realistic goal-driven conversation? How can persona representations reduce language model variance and improve task accuracy? How should models express uncertainty rather than forced confident answers? Do accurate-looking LLM outputs hide structural failures in learning and reasoning? How faithfully do LLMs reflect their actual reasoning in outputs and explanations? Can AI systems balance emotional competence with factual reliability? Do language models develop causal world models or rely on statistical patterns? What pretraining choices and baseline capability constrain reinforcement learning gains? How can real-time alliance measurement improve therapy outcomes?