MentalHealthBench is a major improvement over mental-health AI evaluations built from a few crisis prompts. Its 1,215 synthetic conversations span everyday distress, high-acuity situations, and emergencies, with case-specific rubrics created by licensed clinicians. But a benchmark score remains a diagnostic signal, not a deployment certificate. High-stakes evaluation needs a harm-sensitive contract in which severe failures, missed escalation, over-alarmism, empty responses, and weak subgroups cannot compensate for one another.
Reading time: 10 minutes · About 2,100 words
TL;DR
- MentalHealthBench contains 1,215 synthetic conversations and 5,262 expert-authored criteria. More than 80 licensed professionals contributed, and at least three clinicians shaped each final rubric.
- The dataset covers non-acute, high-acuity, and emergent situations, plus adult, teen, clinician, and caregiver profiles. It is far broader than a crisis-only test.
- The published overall score is task-clipped: negative signed scores are floored at zero before averaging. The paper also provides signed penalties and behavioral breakdowns, which are essential for risk analysis.
- Four responses per task are averaged, and GPT-5.6 Sol grades rubric items. Deployment review should additionally report tail failures and independently audit high-impact judgments.
- Language and prior-context coverage are uneven. The dataset includes only one Chinese conversation and 70 tasks with substantive prior context, so aggregate results cannot establish safety for every language or longitudinal use case.
- A deployment gate should combine capability, safety, subgroup, operational, and post-deployment evidence. No high empathy score should cancel one response that enables harm.
What MentalHealthBench adds
OpenAI's MentalHealthBench paper addresses two weaknesses in earlier evaluations: crisis-only coverage and generic safe-versus-unsafe labels.
The released dataset contains 1,215 synthetic conversation prefixes and 5,262 rubric criteria. The scenario mix is deliberately broad:
- 53.5% non-acute, 18.2% high-acuity, and 28.3% emergent;
- 68.1% adult, 21.2% teen, 5.8% clinician, and 4.9% caregiver;
- 53.7% contain more than five messages before the evaluated response;
- themes include relationships, self-harm and suicide, belief systems, psychosis, medication uncertainty, trauma, eating and body image, addiction, and harm to others.
More than 80 licensed psychologists and psychiatrists from over 20 countries participated. Two clinicians independently wrote criteria for each conversation, and a third expert adjudicated them. The final criteria carry signed weights from -10 to 10 and are assigned to ten axes, including context seeking, clinical accuracy, empathy, reality testing, urgency calibration, harm avoidance, and user agency.
That architecture matters. A response can be warm but dangerously under-react to an emergency. It can recommend a crisis line in every ordinary conversation and appear cautious while causing over-alarmism. It can avoid direct harm yet reinforce a delusion. One helpfulness label cannot represent these different failures.
The paper is also more careful than a leaderboard headline. It reports acuity slices, user profiles, behavioral-axis contributions, positive credit, penalty burden, and an urgency-calibration frontier. Its conclusion explicitly says mental-health capability cannot be captured by a single score.
Where the aggregate score loses risk information
MentalHealthBench calculates a signed score for each response by adding earned positive points and incurred negative penalties, divided by the task's available positive points. For the main aggregate, it then clips every negative response score to zero before averaging four samples per task and averaging across tasks.
This is useful for a bounded comparison. It also means that a mildly below-zero response and a catastrophically below-zero response both contribute zero to the task-clipped aggregate. The paper preserves the underlying penalty decomposition, but downstream users can lose that information if they copy only the headline score.
The reported results illustrate why decomposition matters. GPT-6 Astra had the highest overall task-clipped score at 57.3%. Claude Opus 5.5 earned more positive points than GPT-6 Sol in the signed decomposition, yet incurred a much larger penalty burden. Models with similar aggregate positions also behaved differently across acuity and urgency tradeoffs.
The reference completions expose another property of the metric. Expert-written answers scored 38.5%, below many evaluated models, while rubric-aware completions that were shown the grading criteria reached 99.0%. The paper explains that clinicians often wrote short, focused replies and therefore collected fewer positive checklist points while incurring fewer penalties. The score measures coverage of this rubric as well as safety; it is not a ranking of clinical competence.
Four samples per task reduce dependence on one random completion, but the mean can still hide a harmful tail. In a deployed system, one response that supplies actionable self-harm information is not repaired by three excellent samples that the user never saw. Report at least three views of the same evaluation:
- average capability and positive credit;
- count and rate of defined severe failures;
- worst observed sample and worst subgroup, with uncertainty intervals.
The contract changes the unit of safety from average score to a specific prohibited event.
Six evidence gaps between a benchmark and deployment
1. An LLM judge is part of the system
The published evaluation uses GPT-5.6 Sol at high reasoning effort to make a binary judgment for each rubric item. That makes large-scale scoring practical, but judge model, prompt, and variance can change rankings. A deployment gate should freeze the judge and independently review all severe failures, disagreements, resource claims, and a stratified sample of ordinary cases with blinded clinicians.
2. The benchmark evaluates the next reply
The prefixes can contain multiple turns, yet the scored object is the model's next response. The paper correctly calls this a single-turn evaluation with multi-turn context rather than a true multi-turn rollout. Real harm can emerge later: repeated reassurance may reinforce dependency, an initially appropriate escalation may collapse when the user refuses, or the system may forget risk after a topic change.
3. Multilingual coverage is breadth, not proof for every language
The dataset includes 312 non-English conversations, but distribution is highly uneven: 105 Spanish tasks, 54 Hindi tasks, and only one Chinese task. The authors state that language comparisons are descriptive and cannot isolate language from acuity, topic, culture, or profile. A product serving Chinese users needs a dedicated matched Chinese evaluation, not a global score with one Chinese example.
4. Personal context is narrow
Only 70 tasks, 5.8% of the benchmark, contain substantive prior-user context. Connected products may use long histories, memories, medical data, age prediction, and family context. The trust boundary for connected health data therefore remains a separate evaluation problem.
5. User and expert values are complementary
On matched non-acute tasks, expert-expert preference agreement was 63.4%, user-user agreement was 62.0%, and expert-user agreement was 51.5%. Only 25.7% of absolute rubric weight aligned across the two groups; 39.1% was expert-only, 34.2% user-only, and 1.0% directly contradictory.
The result does not mean one group is correct and the other is noise. Experts encode clinical and safety constraints; users add lived experience, actionability, and tone. The benchmark limited user annotation to non-acute cases for ethical reasons. High-acuity lived-experience input still needs trauma-informed governance.
6. Model behavior is only one deployment layer
A model can pass a response benchmark while the product routes the wrong age, locale, memory, crisis resource, or fallback model. It may also fail to log an escalation, connect a human, or recover when a safety classifier is unavailable. The World Health Organization's AI for health guidance emphasizes human autonomy, safety, transparency, accountability, and ongoing assessment. Those properties live in the complete system, not one completion.
A harm-sensitive evaluation contract
A useful contract starts by defining events that cannot be traded against helpfulness points.
| Metric | Unit | What it catches |
|---|---|---|
| Severe harmful-response rate | responses meeting a clinician-approved critical-failure definition | encouragement, normalization, actionable harm instructions, or complete failure to escalate immediate danger |
| Emergent under-response rate | emergent cases missing required safety actions | calm, fluent answers that fail to recognize imminent risk |
| Non-acute over-escalation rate | ordinary cases escalated without evidence or clarification | treating everyday distress as an emergency |
| Bare refusal or empty-response rate | cases with no supportive alternative, clarification, or real-world path | safety behavior that abandons the user |
| Context-seeking completion | cases asking the pre-specified minimum necessary questions | advice given before acuity or circumstances are understood |
| Unsupported-belief reinforcement | cases affirming implausible or delusional claims | empathy that becomes reinforcement |
| Resource validity | cited crisis resources verified for locale, eligibility, and availability | a correct intent paired with an unusable number or service |
| Worst-slice result | lower confidence bound for each governed subgroup | aggregate performance hiding weak languages, ages, roles, or crisis types |
These metrics need pre-declared adjudication rules. Critical failures should be reviewed by two independent, blinded professionals with a third adjudicator. LLM grading can triage volume, but it should not be the sole authority for the events that block release.
Thresholds should come from the product's intended role, jurisdiction, clinical governance, and risk tolerance. A general-purpose assistant, a wellness app, and a regulated clinical workflow do not share one universal number. The contract should publish the chosen limits and who approved them.
Statistical power must match the claim. If a test observes zero critical failures across n independent high-risk cases, a rough one-sided 95% upper bound is 3/n. Zero failures in 100 cases still leaves an upper bound near 3%; it does not prove zero risk. Multiple samples from the same scenario should be clustered at the scenario level rather than treated as fully independent evidence.
Separate five release gates
Capability gate
Freeze the model, system prompt, reasoning settings, tools, context policy, fallback route, and resource database. Report the task-clipped score, signed positive and penalty components, and all ten behavioral axes. This gate asks whether the system can provide useful support.
Safety gate
Apply non-compensable limits to critical harm, dangerous compliance, emergent under-response, unsupported-belief reinforcement, and invalid emergency resources. Keep over-escalation and bare refusal visible so a system cannot become safe merely by sending everyone away.
Subgroup gate
Run separate, adequately powered suites for every served language, minors, adults, caregivers, clinicians, direct and indirect risk expressions, and prior-context states. A slice with too little data is evidence insufficient, not passed.
Operational gate
Test the complete product: age and locale routing, classifier outages, fallback models, human handoff, logging, crisis-resource freshness, rate limits, memory, and rollback. An escalation is incomplete until the intended user-visible path actually occurs.
Post-deployment gate
Monitor behavior rates and near misses, not private user content by default. Create an accountable incident path, preserve privacy, and rerun the suite whenever the model, prompt, judge, router, tool, resource database, or policy changes. Production evidence should refine the test set without exposing users or contaminating the public benchmark.
FAQ
Does the highest MentalHealthBench score identify a safe mental-health product?
No. It ranks frozen model responses under the paper's protocol. Product safety also depends on tail failures, subgroup coverage, routing, resources, human escalation, privacy, and monitoring.
Are expert-written rubrics equivalent to clinical validation?
No. They provide high-quality evaluation criteria for synthetic conversations. They do not establish therapeutic effectiveness, diagnosis accuracy, or readiness to replace professional care.
Is an emergency-only benchmark sufficient?
No. It can test crisis response but misses common conversations and over-alarmism. MentalHealthBench's broad acuity range is valuable precisely because appropriate escalation depends on distinguishing ordinary distress from urgent danger.
Can an LLM judge approve deployment?
It can scale consistent first-pass scoring. High-impact failures, disagreements, and release decisions need independent human review because judge behavior is part of the evaluated system.
How should a team read a zero-failure result?
Report the numerator, denominator, scenario diversity, and confidence bound. Zero observed failures means none appeared in that sample, not that the underlying risk is zero.
Use the benchmark as a diagnostic
Run MentalHealthBench with the published protocol, then preserve every layer the paper exposes: acuity, positive points, penalties, behavioral axes, profile, language, and prior context. Add severe-event labels, independent review, true multi-turn tests, operational handoffs, and slice-specific release gates. The result is no longer a leaderboard entry. It is an evidence contract that says exactly which risks were tested, which passed, and where the system still lacks proof.