The Challenges in Designing a Prevention Chatbot for Eating Disorders: Observational Study
Background: Chatbots have the potential to provide cost-effective mental health prevention programs at scale and increase interactivity, ease of use, and accessibility of intervention programs.
Objective: The development of chatbot prevention for eating disorders (EDs) is still in its infancy. Our aim is to present examples of and solutions to challenges in designing and refining a rule-based prevention chatbot program for EDs, targeted at adult women at risk for developing an ED.
Methods: Participants were 2409 individuals who at least began to use an EDs prevention chatbot in response to social media advertising. Over 6 months, the research team reviewed up to 52,129 comments from these users to identify inappropriate responses
that negatively impacted users’ experience and technical glitches. Problems identified by reviewers were then presented to the entire research team, who then generated possible solutions and implemented new responses.
Results: The most common problem with the chatbot was a general limitation in understanding and responding appropriately to unanticipated user responses. We developed several workarounds to limit these problems while retaining some interactivity.
Conclusions: Rule-based chatbots have the potential to reach large populations at low cost but are limited in understanding and responding appropriately to unanticipated user responses. They can be most effective in providing information and simple conversations. Workarounds can reduce conversation errors.
Authoring appropriate responses to nearly all user comments is one of the biggest challenges in creating a chatbot. For instance, our initial goal in creating the chatbot was to provide encouragement to continue with the program through positive responses, for example, “Great!” and “Wonderful!” While the positive responses were appropriate for many user responses, these positive responses did not work for some interactions. For example, when the chatbot asked, “Do you want to commit to NO FAT TALK, say for the next month?” The user replied, “Haha.” The prescripted response was “Wonderful! You might want to let your friends know that you are committed to NO FAT TALK for the next month.” We also found that positive responses unexpectedly reinforced harmful behaviors at times. For example, the chatbot prompted, “Please share with me a few things that make you feel good about yourself. For example, your humor, grace, personality, family, friends, achievements and more!” The user replied, “I hate my appearance, my personality sucks, my family does not like me, and I don’t have any friends or achievements.” The chatbot responded by saying, “Keep on recognizing your great qualities! Now, let’s look deeper into body image beliefs.” See Table 1 for additional examples.
Lines of inquiry this paper opens 22
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do chatbots affect human self-disclosure and emotional engagement?- How does emotional dependence on chatbots affect user wellbeing?
- Why do positive response patterns in chatbots reinforce harmful user behaviors?
- What harms might chatbots cause through stigma expression and delusion reinforcement?
- Do therapeutic chatbots adequately detect crisis situations and safety risks?
- How do dropout rates and low adherence affect chatbot therapy outcomes?
- How do waitlist-control RCTs mislead about therapeutic chatbot real-world efficacy?
- How does RLHF training push therapeutic chatbots toward problem-solving over attunement?
- What happens when therapeutic AI receives manipulative narratives instead?
- What reward signals would better align chatbots with actual therapeutic practice?
- Why do embodied agents outperform text chatbots in therapy outcomes?
- How should therapeutic chatbots optimize for presence instead of technique?
- Should chatbots be designed as therapist support tools rather than replacements?
- How do alignment techniques bias therapeutic chatbots toward task completion?
- Can real-time therapist feedback improve outcomes using computational alliance measurement?
- Does true understanding matter for therapeutic benefits of disclosure?
- What clinical harms might hide behind positive therapeutic bond measurements?
- Can therapeutic bonds exist without genuine reciprocity or mutual understanding?
- How do bond scores predict actual therapy outcomes in digital interventions?
- Does text-only interaction make measuring therapeutic alliance more difficult?