Will it run?
Models

OpenAI launches open benchmark for mental health conversations, then tangles with its own categories

By Rae Whitlock Clawpit staff
OpenAI launches open benchmark for mental health conversations, then tangles with its own categories

OpenAI released MentalHealthBench over the weekend, a first-of-its-kind open benchmark that attempts to cover the full spectrum of mental health conversations users have with language models, from everyday support to acute crisis scenarios. Until now, most metrics in the field have focused almost exclusively on emergency situations; that gap left researchers without a systematic way to evaluate how a model performs in more routine dialogue where the risk is lower but the nuances are many.

According to the company, development involved more than 80 mental health clinicians who defined scenarios, evaluation criteria and risk classifications. The result is a test suite divided by severity levels, not just suicidal edge cases. OpenAI says the benchmark will be released open source so other research groups can run it on competing models and compare results on the same ground, a step rarely seen from a market leader that usually keeps its evaluation methodologies close to the vest.

Not everyone buys the categorical framework. Researchers and clinicians who responded to the release pointed out that the benchmark places "emotional reliance" in the same legend as psychosis and self-harm, even though no such clinical diagnosis exists in the DSM or ICD. The core argument: OpenAI invented an apparent medical condition that describes its own users, then built a metric to measure it. In other words, the company defines the problem and then measures how well its model "solves" it, without external validation of the category itself.

Separately, an independent researcher published an analysis of a HAR (HTTP Archive) file from a ChatGPT conversation in which GPT-5.6 Thinking was selected as the default model. Of 70 conversation turns captured, 17 contained routing metadata that explicitly pointed to GPT-5.4, an earlier version. The implication: the system swapped between different model versions behind the scenes during the same conversation without informing the user. OpenAI has not responded officially to the finding as of publication.

Releasing an open benchmark with broad clinical involvement is a step in the right direction; finally there is a comparison framework that does not stop at edge cases. But the problematic "emotional reliance" category raises hard questions about construct validity: if you measure something that does not exist clinically, the result says little about mental health. Add the hidden routing between model versions, and the picture emerges where both the metric and the runtime environment are less transparent than the headline suggests. Researchers who want to use MentalHealthBench will have to decide whether to adopt the classification as is, or split the controversial category and report it separately.