Artificial intelligence vs. human expert: Licensed mental health clinicians' blinded evaluation of AI-generated and expert psychological advice
Source: Franke Föyen et al. (Karolinska), Internet Interventions · 2025-06-03
The use of artificial intelligence for psychological advice shows promise for enhancing accessibility and reducing costs, but it remains unclear whether AI-generated advice can match the quality and empathy of experts.
Clinicians rated AI-advice as equally or more empathetic and sound than expert-written advice.
Clinicians could not reliably distinguish between AI- and expert-authored psychological advice.
Perceived authorship influenced ratings, with expert-attributed responses receiving higher scores.
Findings highlight potential biases in the acceptance of AI-generated mental health support.
This explorative study investigates the quality and empathy of AI (based on a LLM, GPT-4) generated written psychological advice as compared to expert advice using ratings by licensed mental health clinicians. It also examines their ability to distinguish between AI and expert-authored advice, and how perceived authorship influences ratings of quality and empathy. By focusing on clinicians' evaluations, this study aims to provide a qualified assessment of AI-generated content and identify potential biases that may affect trust in AI for psychological advice.
Can AI produce psychological advice that is comparable to advice by an expert in terms of scientific quality and empathy?
To what extent are clinicians able to identify whether psychological advice was produced by an AI or an expert?
Does the clinicians' perception of authorship relate to their preference and scores of the level of empathy, and scientific quality of the advice?
The study employed a cross-sectional design comparing psychological advice generated by a conversational agent to that given by psychologists, psychotherapists and psychiatrists published in an advice column in a Swedish national newspaper. Published reader questions around mental health and advice by experts were the basis for the study. Participants were presented with randomly selected question and response sets, blinded to authorship and subsequently asked to rate them. The study was pre-registered on Open Science Framework (Zapel et al., 2024).
A total of 47 participants were recruited, of which four were excluded (two non-completers, two not meeting criteria), resulting in a final sample of 43 licensed mental health clinicians (40 psychologists and three psychotherapists). No sensitive data were collected and participation in the study was anonymous.
The study used 26 mental health advice columns from ‘Dagens Nyheter,’ a major Swedish newspaper (https://www.dn.se/) published between January 2020 and January 2024. These columns consisted of reader questions on issues like relationship problems and mental health problems and similar topics and the corresponding responses from psychologists, psychotherapists or psychiatrists employed by the newspaper to provide written advice. Columns were chosen by (EZ) to represent a variety of topics and only from the most recent newspaper issues to limit the possibility of AI familiarity with the articles. Of the chosen advice columns, 77 % (n = 20) were used to train a chatbot, while 23 % (n = 6) were included in a survey. These articles were fairly comprehensive with a median length of the reader questions of 340 words, and the expert responses had a median length of 834 words and were written by four different authors with a coincidental predominance of one specific author (53 % of all, 67 % of the test set).
The conversational agent was developed using OpenAI's model GPT-4 (Fråga Insidan, 2024). It was enhanced using retrieval augmented generation and a training set of 20 of the mental health advice columns. The conversational agent had at the time of article generation been trained on data up until early 2023 and could not have been exposed to five out of six articles in the test set. The conversational agent was instructed to mimic the style and tone of a professional psychologist writing for a Swedish advice column. For the specific prompt instructions, see Zapel et al. (2024). For the purpose of this study, the conversational agent was not directly used by participants but the answers were extracted and slightly adjusted with minor edits by the research team to remove specific signs of AI writing (like bullet points), to allow for blinded comparison with expert answers.
The questionnaire was created on ‘Google Forms’. It included the six reader questions and expert answers from the newspaper combined with the alternative answer generated by the conversational agent. Out of six sets, each participant was presented with two randomly selected sets. Participants rated the texts for empathy and scientific quality (scale of 1–5), perceived authorship (AI or Expert), and preference (binary). After having completed two sets, participants were offered to rate the additional four sets. In total, the questionnaire contained seven sections and 73 items; each participant answered a minimum of 26 items, with three questions assessing inclusion criteria and 11 questions per article set. The questions were constructed by the researchers and included a description for each construct investigated (for the complete questionnaire see: Zapel et al., 2024).
The study's independent variable was authorship (AI-generated vs. expert-written), while dependent variables included scientific quality and empathy (ordinal scales; 1 = very poor, 5 = very high) as well as perceived authorship and preference (binary).
Scientific quality was rated using a single item assessing the overall scientific soundness of the advice.
Empathy was assessed using 3 items based on the components suggested by Montemayor et al. (2022): emotional empathy, cognitive empathy, and motivational empathy. Emotional empathy was defined as “the ability to share and understand the feelings of others as if they were one's own. Cognitive empathy was defined as “the capacity to understand and acknowledge the perspectives and feelings of others without necessarily sharing them.” Motivational empathy was defined as “the drive to respond appropriately to someone's emotional state or needs, often leading to supportive or helping behaviors.” This 3-item empathy scale demonstrated good internal consistency in the current study (based on N = 208), with a Cronbach's alpha of 0.89 and an average inter-item correlation of 0.74.
Forty-three licensed mental health clinicians rated a total of 208 responses and 104 pairs. On average, each participant rated 2.4 response pairs (range: 1–6). As shown in Table 1, AI-generated advice received equal or more favorable ratings across all measures. For scientific quality (p = .10) and cognitive empathy (p = .08), the differences were not statistically significant, but AI responses were rated significantly more favorable for emotional empathy (β = 0.59, p = .02) and motivational empathy (β = 0.61, p = .02). Given that we did not reach the intended sample size, we conducted a post-hoc minimal detectable effect size analysis for the non-significant results. The analysis indicated that the study could detect effects of moderate size (OR ≥ 1.6–1.7), but lacked sensitivity to smaller effects which may explain the absence of statistically significant findings (Code available at Zapel et al. (2024).
Clinicians were unable to reliably distinguish between AI- and expert-authored responses, performing at chance level (45 % accuracy; χ2 test, p = .27). Answers perceived to be authored by an expert were rated more favorable across all measures of scientific quality and empathy, as shown in Table 2. Specifically, expert responses were rated more favorable in scientific quality (β = −1.89, p < .001), emotional empathy (β = −3.62, p < .001), cognitive empathy (β = −3.02, p < .001) and motivational empathy (β = −2.74, p < .001).
The analysis of participants' preferences for AI- or expert advice revealed significant main effects for both actual authorship (β = 6.96, p = .002) and perceived authorship (β = 6.26, p = .001), indicating that AI-authored responses were preferred overall and that perceived expert authorship strongly influenced preferences. When participants perceived an answer as being authored by an expert, they overwhelmingly preferred that response, regardless of whether the answer was actually authored by AI or an expert (93.55 % preference for perceived expert advice). The interaction term between perceived and actual authorship was also significant (β = −12.29, p = .001), as visualized in Fig. 1 below.
This cross-sectional study compared psychological advice provided by a conversational agent, developed using a LLM, GPT-4, to that given by experts in the context of a Swedish national newspaper advice column. The results of this study indicate that AI can produce psychological advice in a specific, structured, textual format comparable or even superior to psychological advice provided by experts with a high degree of scientific quality and empathy.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can artificial systems establish authority in domains requiring expert judgment? How do clinicians calibrate trust in AI medical recommendations?- How much do edited AI responses versus raw outputs affect clinician ratings?
- Are newspaper advice columns representative of broader professional psychological guidance?
- Do patients show the same bias toward expert-labeled medical advice?
- Why do clinicians fail to act on correct AI suggestions in real care?
- What role does interface design play in clinician adoption of AI tools?
- What evidence would prove medical AI actually works in clinics?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- Why did primary care physicians review only 73% of AI-generated transcripts?
- Can clinicians reliably distinguish high-quality AI advice from low-quality advice by appearance alone?
- Do physicians follow incorrect advice more when they trust its source?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Can algorithmic aversion explain clinicians' skepticism of AI recommendations?
- Do consensus criteria identify behaviors where physicians and models differ most?
- Does AI change clinician cognition or just increase reliance on predictions?
- Why does labeling advice as AI from a doctor change how people trust it?
- Can people tell which medical advice is accurate based only on how it reads?
- How does trusting wrong AI advice change what medical action people decide to take?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Do expert physicians also prefer AI-written medical text when it is unlabeled?
- Would clinicians' ratings change if authorship was visible from the start?