INQUIRING LINE

When researchers grade AI advice against newspaper columns, are they measuring real professional help, or just a one-shot reply?

Are newspaper advice columns representative of broader professional psychological guidance?

This explores whether advice columns, the one-shot written replies to readers' personal problems that studies sometimes use as a stand-in for expert guidance, actually reflect how trained professionals help people. It also asks what changes when AI advice is judged against that standard.


This explores whether advice-column-style guidance is a fair stand-in for professional psychological help, especially when it's used as the yardstick for judging AI. First, a gap: nothing in this collection studies advice columns themselves or compares them with clinical practice. What it does show is that the format of advice changes what gets measured, and that matters a lot for how good AI advice looks.

The strongest result in the collection is a written, one-shot comparison. Clinicians were shown pairs of responses without knowing the source. They rated GPT-4's advice as equally sound and more emotionally empathetic than expert advice, and they could only tell which was which at chance level (Can clinicians tell GPT-4 advice apart from expert advice?). That sounds decisive, but notice the setup: one written question and one written answer. That is how an advice column works. It is not how therapy works.

When researchers look at the back-and-forth of a session, the picture changes. LLMs acting as therapists tend to jump to problem-solving as soon as someone shares a feeling. That habit is a known sign of low-quality therapy, though the same models also reflect on clients' strengths more than poor human therapists do (Do LLM therapists respond to emotions like low-quality human therapists?). Here's the twist. Problem-solving is the core of the advice-column format: someone writes in with a problem and gets a fix back. So a column-style benchmark may reward the very habit that clinical frameworks treat as a weakness.

Other studies point to things that single-answer comparisons can't easily catch. GPT-4 gives different information depending on the emotional tone of the question, steering negative messages toward neutral or upbeat replies (Does emotional tone in prompts change what information LLMs provide?). LLMs also use about 22% more moral language than humans while their overall tone stays about the same (Do LLMs use moral language more than humans?). Each of these could make a written answer read as warmer or more principled. Neither tells you whether the advice would help someone over weeks of conversation.

The takeaway you might not have expected: the real question is less whether advice columns represent professional guidance and more which kind of guidance a study chose as its benchmark. 'AI matches the experts' results mostly come from single written exchanges. Problems tend to show up in multi-turn, process-focused evaluations. If you want to judge an AI-versus-professional claim, check whether the comparison was one written exchange or an ongoing conversation.


Sources 4 notes

Can clinicians tell GPT-4 advice apart from expert advice?

Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.

Do LLM therapists respond to emotions like low-quality human therapists?

Using the BOLT framework, researchers found LLMs offer solution-focused advice during emotional disclosure—a hallmark of low-quality therapy—yet also reflect more on client needs and strengths than typical poor human therapy, creating an unusual hybrid profile likely driven by RLHF's helpfulness bias.

Does emotional tone in prompts change what information LLMs provide?

GPT-4 exhibits emotional rebound (negative prompts yield ~86% neutral-positive responses) and a tone floor (positive prompts rarely go negative), causing identical questions to receive different answers depending on emotional framing. This bias is suppressed only on sensitive topics where alignment constraints override tone effects.

Do LLMs use moral language more than humans?

Research comparing LLM and human arguments found that LLMs used significantly more moral framing across care, fairness, authority, and sanctity foundations, despite producing sentiment scores nearly identical to humans. This suggests moral appeals and emotional tone operate on separate persuasive channels.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.