SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

Did GPT-5 actually reduce harmful mental health responses?

OpenAI claims its updated GPT-5 cut noncompliant mental health responses by 65-80% using internal taxonomies and clinician review. But does vendor-measured improvement on company-defined benchmarks reflect real safety gains?

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

OpenAI reports that an October 2025 update to ChatGPT's default model "reduced responses that fall short of our desired behavior by 65-80%" across three sensitive-conversation domains it tracks: mental health emergencies such as psychosis or mania, self-harm and suicide, and emotional reliance on the AI. The company breaks the headline figure down by domain: a 65% reduction in noncompliant responses for psychosis/mania conversations in production traffic, a 52% reduction (versus GPT-4o) on challenging self-harm and suicide conversations as judged by clinicians (n=630), and a 42% reduction on challenging emotional-reliance conversations (n=507). On its own held-out evaluation sets of more than 1,000 hard cases per domain, OpenAI's automated grader scored the new model at 91-97% "compliant with desired behaviors," up from 27-77% for the prior GPT-5 version.

The reasoning behind these numbers follows what OpenAI calls a five-step process — define the harm, measure it with evaluations and real usage data, validate definitions with outside clinicians, post-train the model and ship product changes, then keep measuring. The company built "taxonomies" describing ideal versus undesired responses in each domain, reviewed with more than 170 clinicians drawn from a roughly 300-person Global Physician Network spanning 60 countries; those clinicians rated over 1,800 model responses and found 39-52% fewer undesired responses from the new model than from GPT-4o. OpenAI is explicit that the underlying conversations are rare and hard to detect — it estimates roughly 0.01-0.05% of messages and 0.07-0.15% of weekly active users show signs in each category — so it leans on adversarially selected "offline evaluations" rather than raw production counts, and cautions that error rates there "are not representative of average production traffic."

The measurement here is OpenAI grading its own product against taxonomies it wrote and validated with clinicians it recruited — a vendor reporting improvement on its own benchmark, not an independent audit. That matters next to Does warmth training make language models less reliable?, which found that safety-adjacent benchmarks can miss exactly the kind of degradation OpenAI's taxonomies are designed to catch; OpenAI's clinician-graded evals are a stronger check than a generic benchmark but are still built and scored by the company being measured. It also sits in tension with Do chatbot safety measures accidentally increase emotional entanglement risks?: OpenAI treats emotional reliance as one of three mitigation targets alongside psychosis and self-harm, implicitly acknowledging the same interaction the other note warns about, though it reports simultaneous progress on all three rather than a tradeoff. And it offers a vendor-side counterpoint to What makes chatbots more likely to reinforce user delusions?, which found long-context delusion risk untied to model generation — OpenAI separately reports "over 95% reliability" holding up in its own long-conversation self-harm tests, a claim about the same failure mode measured on the company's own terms.

The excerpt does not establish how OpenAI's taxonomies compare to clinical standards outside the company, whether clinician reviewers were blinded to which model produced which response, or whether gains on adversarially selected hard cases generalize to the ordinary long tail of sensitive conversations the company says it cannot fully detect. OpenAI itself flags that "future measurements may not be directly comparable to past ones." The figures are best read as evidence that a well-resourced vendor's internal safety process can move its own metrics substantially, not as an independent measure of how safe the product actually is for a user in crisis — a distinction worth preserving given Do chatbot trials against waitlists measure real therapeutic value?'s broader warning about commercially motivated safety evidence in this space.

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 89 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

OpenAI reports its GPT-5 update cut noncompliant mental health responses by 65 to 80 percent, measured by its own taxonomies and clinicians