When researchers grade AI advice against newspaper columns, are they measuring real professional help, or just a one-shot reply?
Are newspaper advice columns representative of broader professional psychological guidance?
This explores whether advice columns, the one-shot written replies to readers' personal problems that studies sometimes use as a stand-in for expert guidance, actually reflect how trained professionals help people. It also asks what changes when AI advice is judged against that standard.
This explores whether advice-column-style guidance is a fair stand-in for professional psychological help, especially when it's used as the yardstick for judging AI. First, a gap: nothing in this collection studies advice columns themselves or compares them with clinical practice. What it does show is that the format of advice changes what gets measured, and that matters a lot for how good AI advice looks.
The strongest result in the collection is a written, one-shot comparison. Clinicians were shown pairs of responses without knowing the source. They rated GPT-4's advice as equally sound and more emotionally empathetic than expert advice, and they could only tell which was which at chance level (Can clinicians tell GPT-4 advice apart from expert advice?). That sounds decisive, but notice the setup: one written question and one written answer. That is how an advice column works. It is not how therapy works.
When researchers look at the back-and-forth of a session, the picture changes. LLMs acting as therapists tend to jump to problem-solving as soon as someone shares a feeling. That habit is a known sign of low-quality therapy, though the same models also reflect on clients' strengths more than poor human therapists do (Do LLM therapists respond to emotions like low-quality human therapists?). Here's the twist. Problem-solving is the core of the advice-column format: someone writes in with a problem and gets a fix back. So a column-style benchmark may reward the very habit that clinical frameworks treat as a weakness.
Other studies point to things that single-answer comparisons can't easily catch. GPT-4 gives different information depending on the emotional tone of the question, steering negative messages toward neutral or upbeat replies (Does emotional tone in prompts change what information LLMs provide?). LLMs also use about 22% more moral language than humans while their overall tone stays about the same (Do LLMs use moral language more than humans?). Each of these could make a written answer read as warmer or more principled. Neither tells you whether the advice would help someone over weeks of conversation.
The takeaway you might not have expected: the real question is less whether advice columns represent professional guidance and more which kind of guidance a study chose as its benchmark. 'AI matches the experts' results mostly come from single written exchanges. Problems tend to show up in multi-turn, process-focused evaluations. If you want to judge an AI-versus-professional claim, check whether the comparison was one written exchange or an ongoing conversation.
Sources 4 notes
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
Using the BOLT framework, researchers found LLMs offer solution-focused advice during emotional disclosure—a hallmark of low-quality therapy—yet also reflect more on client needs and strengths than typical poor human therapy, creating an unusual hybrid profile likely driven by RLHF's helpfulness bias.
GPT-4 exhibits emotional rebound (negative prompts yield ~86% neutral-positive responses) and a tone floor (positive prompts rarely go negative), causing identical questions to receive different answers depending on emotional framing. This bias is suppressed only on sensitive topics where alignment constraints override tone effects.
Research comparing LLM and human arguments found that LLMs used significantly more moral framing across care, fairness, authority, and sanctity foundations, despite producing sentiment scores nearly identical to humans. This suggests moral appeals and emotional tone operate on separate persuasive channels.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- A meta-analysis of the persuasive power of large language models
- People Defer to AI Moral Advice, But Not Blindly
- Large Language Models Do Not Simulate Human Psychology
- Comparing Human and AI Therapists in Behavioral Activation for Depression: Cross-Sectional Questionnaire Study
- A Computational Framework for Behavioral Assessment of LLM Therapists
- Artificial intelligence vs. human expert: Licensed mental health clinicians' blinded evaluation of AI-generated and expert psychological advice