SDSignal Desk

Introducing MentalHealthBench

Sep 23, 2026, 3:00 AM · OpenAI

Image: OpenAI

OpenAI opens MentalHealthBench with more than 80 licensed clinicians across 22 countries — a shared yardstick for how models handle everyday stress, high-acuity distress, and emergencies.

Why it matters

People already bring relationship conflict, everyday stress, and harder moments to chatbots. OpenAI’s MentalHealthBench, published September 23, 2026, is an open evaluation built with more than 80 licensed psychologists and psychiatrists from 22 countries, spanning 19 languages and nearly 20 subspecialties.

Most prior evals leaned on emergencies and coarse pass/fail rules. This one grades realistic conversations across non-acute, high-acuity, and emergency scenarios — adults, teens 13–17, caregivers, and clinicians — against expert rubrics that reward helpful behaviors and penalize harmful ones. With ChatGPT claiming more than a billion weekly users, how models handle this continuum is no longer a niche research question.

From the desk

We’re reading this as useful AI doing the hard, unglamorous work: measuring care, not just capability.

The method matters. Experts write weighted criteria for each synthetic conversation’s last turn — positive points for asking the right question or preserving agency, negative for overconfident advice or guessing feelings. At least three clinicians review each case; criteria stick only with multi-expert agreement. An automated grader (GPT-5.6 Sol) scores responses against those rubrics. That’s closer to clinical judgment than a single toxicity filter, and OpenAI is releasing the benchmark so others can rerun and stress-test it.

Useful AI earns the benefit of the doubt when it invites outside scrutiny and routes people toward real help — localized crisis lines, Trusted Contact, ChatGPT for Teens protections — instead of pretending the model is therapy. OpenAI is explicit: ChatGPT is not a substitute for professional care. We agree. The upside is models that seek context, keep user agency intact, and nudge toward trusted humans when things escalate.

The harm side is real. Synthetic conversations, even privacy-preserving ones, are still proxies. A teen persona flagged via system message may miss product-specific youth safeguards. Rubrics reflect expert consensus, while a separate 44-user study found people value tone and practical next steps more than clinicians emphasize — and experts care more about gathering context. If labs optimize only for the expert scoreboard, responses can feel cold or incomplete to the people actually typing at 2 a.m.

I’m watching whether independent labs adopt MentalHealthBench as a shared standard, whether score gains translate into fewer harmful ChatGPT incidents, and whether teen-path and emergency routing keep improving as models get more fluent.

Context

OpenAI frames MentalHealthBench as a successor to HealthBench and HealthBench Professional, co-created with a global clinician cohort including APA CEO Dr. Arthur Evans among the quoted voices. Companion work includes research grants, Partnership on AI convenings, and nods to independent efforts such as Transluce’s mental health evaluation.

Who feels it

Model labs and eval teams
An open, rubric-scored mental-health suite that spans acuity levels and personas — and invites outside reruns — raises the bar beyond emergency-only checks.
Clinicians and safety researchers
Weighted expert criteria and multi-reviewer consensus give a shared language for what “good” looks like in ambiguous, non-emergency support chats.
Product and trust teams
User-vs-expert gap on tone and actionable steps is a product design signal, not just a research footnote.
Parents and teen-product stewards
Youth scenarios were clinician-reviewed, but system-message teen personas may not mirror every in-product safeguard — verify before claiming coverage.

What to watch

  1. Independent labs publishing MentalHealthBench scores alongside their own safety cards.
  2. Whether ChatGPT incident rates in distress conversations move as models climb the benchmark.
  3. How teen-path and Trusted Contact / crisis-routing features evolve with the next model drops.
  4. Whether user-valued tone and practical next steps get folded into future rubric versions without diluting clinical caution.

Read the original

Continue at the source.

OpenAI