This past week, OpenAI released Mental Health Bench, their benchmark for how generative AI systems should handle questions around mental health. The announcement was accompanied by a public post from OpenAI and a scientific paper on the specifics of their benchmark. But the benchmark has serious problems that should concern physicians whose patients are already using it.
What They Created
The benchmark is a compilation of 1,215 synthetic mental-health conversations, including:
Non-acute everyday stress (53.5%)
High-acuity distress (18.2%)
Emergencies (28.3%)
It includes questions from adults, teens 13–17, caregivers, and clinicians, in multiple languages. According to the paper, over 80 psychiatrists and psychologists from 22 countries participated. They authored 5,262 rubric criteria, weighted on a scale of -10 to +10, across 10 dimensions in behavioral health, like safety, user agency, actionable guidance, clinical accuracy, empathy, and more. The entire benchmark is open, which means you can see all the questions and the rubric criteria written to guide the answers here.
That’s a lot of numbers. So let’s pause one second to talk about what they mean. 1,215 questions made up of conversations (turns) between a patient and an AI. They were modeled based on real scenarios asked to ChatGPT. The scenarios were stripped of identifiers and then fed to an LLM to create synthetic versions of the conversations.
The conversations included multiple rounds (turns ) between the patient and an AI. Despite that, only the last response or statement was scored by the clinician. That seems a little artificial, and OpenAI acknowledges that in their post. Relationships between patients and their mental health clinicians are long-term. This benchmark is grading the ideal answer to the final response in the conversation.
One more thing about the benchmark. Even though it’s based on real encounters between humans and ChatGPT, OpenAI says the scenario mix “does not represent how often these topics occur in ChatGPT”. So, it requires some trust in OpenAI that their LLM is able to identify common mental health complaints, strip away non-relevant details, and develop its own sample cases.
Were Patient Opinions Included?
In the scientific paper, OpenAI says that they surveyed 44 patients from 16 countries. The purpose was to ask patients about what they thought was the most helpful response. Interestingly, the patients ranked “warm tone and practical next steps” highest. However, clinicians ranked “gathering context first, interpreting ambiguity carefully, safety behaviors” as most important.
This is something that clinicians have seen in practice for decades. The “best” response differs depending on whether you are the patient or the doctor. Clinicians are far more focused on safety. How did OpenAI resolve the tension? They documented the patient preferences, but did not apply them. Only the rubric criteria written by the clinicians were used to train the model.
So, the patient opinions became an interesting footnote in the study. But they did not play a role in grading.
Which AI Models Were Tested?
The scientific paper details more than just the creation of the benchmark. OpenAI tested 17 different models, 7 of them their own.
The top 5 spots went to:
GPT-6 Astra, 57.3%
GPT-6 Sol, 53.9%
Claude Opus 5.5, 52.4%
GPT-6 Luna, 50.2%
Muse Spark 1.3, 48.6%
There are, however, some important things to note about the methods and findings.
First, they published task-clipped scores. That requires a little explanation. When the experts wrote rubric criteria on a scale of -10 to +10, they scored beneficial behaviors positively and negative behaviors (like harmful statements) negatively. But a task-clipped score sets a floor of zero for negative scores.
What does that mean practically? Instead of adding up all the positive and negative scores to get a final number that could fall below zero, task-clipping prevents that. It creates a score that is less susceptible to outliers. For example, take a model that performs fairly on most questions but then has a category where it has dangerous answers and severely negative scores. Instead of reporting an overall score of -12%, the effect of those severely negative scores is blunted. By task-clipping, you ensure a positive final percentage. It doesn’t mean the model performed well on those dangerous questions, but it does lessen the effect of catastrophic failure in specific categories.
If you’re not familiar with such statistical methods, that’s ok. I wasn’t either. But it’s a sticking point for us in medicine. Answers are not just right or wrong; they can be catastrophic and result in real harm. Those have to be scored negatively. And statistics like task-clipping reduce the impact that those negative scores can have.
The OpenAI paper reported task-clipped scores. When you recalculate using the full range, the top 5 models look like this:
GPT-6 Astra: 50%
GPT-6 Sol: 47%
GPT-6 Luna: 42%
Claude Opus 5.5: 40%
Muse Spark 1.3: 32%
The scores for the top 5 dropped, but none dipped into the negative range. Meaning no model was graded as harmful, even though the paper itself documents that these models produce harmful responses regularly.
Which brings me to the second important finding: these models performed terribly. The 17 models tested averaged -19% to -42% on their responses. To be clear: that's not a small slice of bad performance. It means up to half of what these models generate on mental health questions are answers that clinicians scored as actively harmful. They don't just miss the mark on some questions. They produce dangerous responses with regularity.
What Other Quirks Are There?
The entire process is to establish a benchmark against which we can all agree these models should be tested. But the process has some serious issues:
Each of the 17 models was graded by an AI model, GPT-5.6 Sol from OpenAI. So an LLM sits as judge over other LLMs, a process which has been shown to affect scores because LLMs have biases and prefer themselves over other models. The authors acknowledged this limitation.
The experts were paid by OpenAI, the benchmark was created and administered by OpenAI, the paper was published by OpenAI, and no peer review was performed.
The benchmark was published the same week as NPR’s review of legal records showing at least 75 lawsuits had been filed in federal and state courts against AI developers over alleged harms from chatbots, many involving children and including several deaths by suicide.
What This Means For You
Your patients are already doing this. The paper’s background section cites the APA’s 2026 survey: 77% of psychologists said their patients reported using AI, and more than a third said patients were using it as an additional mental-health provider.
So what do you tell them? A benchmark score is not a clinical trial. Fifty-seven percent rubric compliance on synthetic conversations tells you nothing about whether the model helped, harmed, or just sounded caring to a real person in distress. It tells you even less about the product your patient actually used, with its own guardrails, memory, and failure modes.
At most, it’s a diagnostic tool for model developers. Don’t use it as evidence that AI is getting good at mental-health support.
When a company grades its own model with its own grader on its own rubric, and the result is 57%, the right clinical response is the same one you’d give to a drug trial with a 57% response rate and no control arm: interesting, keep studying, don’t prescribe it yet.


