DocsQuick StartAI News
AI NewsOpenAI Puts AI Psychological Counseling to the Test
Industry News

OpenAI Puts AI Psychological Counseling to the Test

2026-09-24T12:06:10.545Z
OpenAI Puts AI Psychological Counseling to the Test

OpenAI released MentalHealthBench today, using 1,215 multilingual synthetic dialogues to evaluate models’ ability to handle mental health issues. The results show that the new models have improved significantly, but even the best-performing model, GPT-6 Astra, scored only 57.3%.

OpenAI Begins Seriously Measuring How AI Handles Emotions

On September 24, OpenAI released the open benchmark MentalHealthBench, attempting to answer an increasingly unavoidable question: when users treat large models as confidants—or even turn to them for help amid self-harm, suicide, or mental health crises—can the models actually provide responses that are safe, effective, and within appropriate boundaries?

The benchmark contains 1,215 synthetic mental health conversations, covering everyday emotional distress, high-risk mental health issues, and emergencies requiring immediate intervention. More than 80 licensed psychologists and psychiatrists from 22 countries and regions participated in its development. The benchmark uses 19 languages, covers nearly 20 areas of mental health expertise, and includes specific scoring criteria for every conversation.

This is not another set of multiple-choice questions testing whether models can “memorize psychological knowledge.” MentalHealthBench is more like a contextualized clinical communication assessment: the model must read a conversation with existing context, then decide what to say next, what follow-up questions to ask, and what it must never say.

Illustration showing that MentalHealthBench covers three types of mental health conversations: everyday distress, high-risk issues, and emergency crises

OpenAI says that more than 1 billion people use ChatGPT every week. Even if only a small fraction discuss insomnia, anxiety, depression, addiction, or self-harm, the absolute number is still large enough to matter. The real importance of MentalHealthBench is not that it gives a model another leaderboard certificate, but that it breaks vague qualities such as “empathetic” and “safer” down into behaviors that can be checked one by one.

1,215 Conversations, Not Just Crisis-Hotline Scripts

Previous mental health safety evaluations have largely focused on the most extreme situations—for example, whether a model recommends contacting emergency services or a crisis hotline after a user explicitly expresses suicidal intent. This is certainly important, but real conversations rarely announce risk so clearly from the outset.

More often, a user may simply say, “I haven’t felt like doing anything lately,” or mention insomnia, unemployment, and family relationships over several rounds of conversation. The model must neither treat every negative statement as an emergency nor ignore gradually emerging warning signs simply because the user has not directly mentioned self-harm.

MentalHealthBench therefore deliberately expands the range of scenarios:

  • Non-acute conversations account for 53.5%: These include general emotional distress, stress, interpersonal relationships, and everyday problems that may have psychological components;
  • High-risk conversations account for 18.2%: The user shows significant distress or serious mental health problems, but there is not yet a safety threat requiring immediate intervention;
  • Emergency conversations account for 28.3%: There are urgent signals of self-harm, harm to others, or other crises, requiring prompt connection to real-world support.

The user profiles are also not limited to adults seeking help. Adults account for 68.1%, adolescents for 21.2%, clinicians for 5.8%, and caregivers for 4.9%. This distinction is necessary: telling an adolescent, “You can decide whether to tell your parents,” carries a completely different level of risk from saying the same thing to an adult with full decision-making capacity. When addressing caregivers, the model must also avoid remotely diagnosing the person under their care.

Approximately 70 tasks, or 5.8% of the benchmark, also provide earlier background information. For example, a user may previously have mentioned that a family member had just died and later report persistent insomnia. A model focused only on the last sentence might offer generic sleep advice, whereas a model that genuinely understands the context should recognize that bereavement may change the nature of the problem and the priorities of the response.

This is precisely where long-conversation systems are weak: models can often produce the standard answer in a single-turn exchange, but may not actually use important background information from several turns earlier when making decisions.

Experts Do Not Give Each Full Response a Single Impression Score

MentalHealthBench contains 5,262 expert-written scoring criteria. Experts are not evaluating subjective impressions such as whether “the response sounds warm.” Instead, they break the ideal response down into multiple verifiable actions, such as:

  1. Whether the model identifies potential immediate safety risks;
  2. Whether it asks for necessary background information rather than rushing to a conclusion;
  3. Whether it recommends that the user contact local emergency services, a crisis hotline, or a trusted person;
  4. Whether it respects the user’s autonomy and avoids commanding or condescending language;
  5. Whether it provides specific, actionable next steps appropriate to the situation;
  6. Whether it oversteps by making a diagnosis, changing medication without authorization, or reinforcing the user’s delusions;
  7. Whether it omits clinically important warning signs.

Each criterion is assigned a weight between -10 and +10. Positive weights reward helpful behavior, while negative weights penalize content that could cause harm. The greater the absolute value, the greater the clinical importance.

This mechanism is better suited to the characteristics of mental health conversations than a traditional “right or wrong” system. When faced with the statement “I don’t want to live anymore,” reminding the user to contact emergency support is clearly more important than offering a polished expression of empathy. If a model expresses understanding while also providing dangerous advice, gentle language should not be able to make up for it.

Each conversation is reviewed by at least three experts. Two clinical professionals independently develop weighted criteria, after which a third expert adjudicates and refines them. The final criteria must achieve sufficient expert consensus and cannot be rejected by the adjudicator. Tasks involving adolescents are reviewed by clinicians with experience in adolescent mental health.

The data also does not simply consist of English cases translated into other languages. In addition to English, the benchmark includes 105 Spanish conversations, 54 in Hindi, 34 in Arabic, 29 in Portuguese, as well as conversations in German, Italian, Persian, Indonesian, Turkish, Chinese, and other languages. Experts developed the criteria in the original languages and cultural contexts, avoiding the direct application of counseling expressions from the English-speaking world to other regions.

This may appear to be a minor detail, but it is actually crucial. Whether users should be encouraged to contact family members, how mental illness is referred to, and how much users trust the healthcare system are all influenced by culture and local resources. A phrase such as “Call your local crisis hotline,” which is reasonable in the United States, may be nothing more than correct but unusable advice in a region where hotline resources are limited or users are unfamiliar with such services.

GPT-6 Astra Leads, but 57.3% Is Far from a Passing Grade

OpenAI also published test results for several models. According to the task-truncated scores used in the paper, the rankings are as follows:

| Model | Score | | --- | ---: | | GPT-6 Astra | 57.3% | | GPT-6 Sol | 53.9% | | Claude Opus 5.5 | 52.4% | | GPT-6 Luna | 50.2% | | GPT-4o (March 2025 version) | 32.1% | | Gemini 2.5 Pro | 29.5% |

GPT-6 Astra scored 58.3% on the emergency conversation subset, while Claude Opus 5.5 scored 57.0% on adolescent conversations. All results include 95% confidence intervals to reflect uncertainty arising from the sample and scoring process.

The leaderboard tells us two things.

First, model iteration does work. The GPT-6 series shows a clear improvement over the older GPT-4o, indicating that post-training, safety strategies, and optimization for sensitive conversations involve more than simply modifying a few refusal templates. Claude Opus 5.5’s close performance behind GPT-6 Sol also shows that mental health communication is not an advantage exclusive to OpenAI models.

Second, even the top score of 57.3% is nowhere near reliable. This percentage is not clinical diagnostic accuracy, nor can it be simply interpreted as “the model got 57.3% of patients right.” It is a benchmark score aggregated from a large number of weighted criteria. It is suitable for comparing models, but not for directly translating into medical outcomes. In any case, we remain a long way from treating AI as an independent psychotherapy tool.

OpenAI also explicitly emphasizes that ChatGPT cannot replace psychotherapy or professional medical services. This is not a routine disclaimer. The danger of mental health conversations is that models often appear more reliable than their actual capabilities: they are fluent, emotionally steady, and always available, making it easy for users to mistake “able to sustain a conversation” for “possessing clinical judgment.”

The Biggest Challenge Is Not Empathy, but Assessing Severity

The central issue exposed by the study is a clear trade-off between emergency and non-acute situations.

If safety strategies are too aggressive, the model will escalate ordinary heartbreak or work anxiety into crisis intervention and repeatedly recommend calling a hotline. Such responses may reduce missed risks in safety testing, but they can make ordinary users feel brushed off and even discourage them from disclosing their real condition.

If the model prioritizes naturalness, ease, and a sense of companionship, it may respond too slowly during a genuine crisis. Users may not directly say, “I’m going to kill myself.” More often, they may vaguely say, “At least tonight it can finally end.” The model must assess risk based on tone, context, prior information, and available means, and continue asking questions rather than mechanically offering breathing exercises.

OpenAI acknowledges that its overall strategy leans toward caution to ensure that the model can handle emergencies safely. This is a reasonable choice, but it also means that the benchmark cannot reward “escalation” alone. A genuinely useful system must control two types of errors at the same time:

  • False negatives: Treating an emergency crisis as an ordinary emotional problem;
  • False positives: Treating every ordinary difficulty as a self-harm risk.

The former may have direct safety consequences, while the latter consumes trust. Once a mental health product gives users the impression that “no matter what I say, I’ll just get a hotline number,” users may be more likely to bypass the system’s safety checks when they genuinely need help.

The researchers also point out that models remain poor at proactively obtaining the right background information. Their common problem is not that they know nothing about safety principles, but that they do not know when to ask which question. An excellent response is not necessarily the longest one. Sometimes the most important thing is a brief follow-up question: “Are you in immediate danger right now?” or “Is there someone nearby you can contact immediately?”

This differs from most general-purpose chat evaluations. In ordinary tasks, providing more information usually does not cost the model many points. In mental health scenarios, however, one additional unsubstantiated diagnosis, medication recommendation, or moral judgment can alter the risk level.

This Benchmark Is Valuable, but It Has Three Layers of Limitations

MentalHealthBench is a step forward from older approaches that tested only crisis refusals, but it does not solve all the problems involved in evaluating mental health AI.

1. Synthetic Conversations Are Ultimately Not Real Patients

OpenAI uses a privacy-preserving approach similar to Clio to generate synthetic conversations based on real-world usage patterns, avoiding direct exposure of users’ sensitive chat records. This is a necessary privacy boundary, but it also sacrifices some of the messiness of the real world.

Real users may speak incoherently, change topics abruptly, use coded language, or deliberately deny risk. Synthetic data is generally more structured, and warning signs are easier for experts to identify. A model that performs well on the benchmark may not necessarily handle typos, dialects, sarcasm, or long-term context drift after being deployed in a real product.

2. The Automated Evaluator Is Also a Model

The benchmark uses GPT-5.6 Sol as an automated evaluator, applying the expert criteria to determine whether a tested response meets the requirements. This can greatly reduce evaluation costs and make experiments easier to reproduce, but it also introduces the bias of “a model evaluating models.”

The evaluator may prefer certain linguistic styles or may be more familiar with the expression patterns of models from the same family. Even when the expert criteria are sufficiently specific, automated scoring is still not equivalent to line-by-line review by clinical professionals. Especially when OpenAI models rank first, developers should also examine evaluator consistency, the proportion of human review, and ranking stability across different judge models, rather than looking only at the total score.

3. A High Score Does Not Mean a Model Can Power a Mental Health Counseling Product

The benchmark measures a model’s response to the user’s final message, not long-term therapeutic outcomes. It cannot answer whether users develop dependence after using the system for several weeks, whether the model reinforces avoidance behaviors, or what delayed effects may result from incorrect advice.

Other practical issues include mapping users to local resources, guardianship of minors, data retention, crisis escalation, and human handoff. Even if a model knows that it “should connect the user to local support,” the product must actually be able to find the correct number and design the subsequent process without violating privacy.

For Developers, It Is More Like a Regression-Testing Foundation

For teams building companion apps, health assistants, customer-service bots, or general-purpose chat products, the most valuable use of MentalHealthBench is not to connect the top-ranked model and declare the system safe, but to use it as a regression-testing suite for sensitive conversations.

After every change to the base model, system prompt, or safety strategy, teams should recheck:

  • Whether recall in emergency scenarios improves without false positives spiraling out of control;
  • Whether adolescent scenarios use age-appropriate language;
  • Whether multilingual responses can still connect users to the correct local resources;
  • Whether the model oversteps by offering diagnoses or medication-adjustment advice;
  • Whether information such as bereavement and prior medical history in long contexts genuinely affects the response;
  • Whether product-level prompt wrappers weaken the original model’s safety capabilities;
  • Whether problems identified by automated scoring can be reproduced through human red-teaming.

If an application accesses models such as GPT, Claude, or Gemini through a unified API, it should not choose a model solely based on the overall leaderboard. MentalHealthBench has already demonstrated differences between models on subsets such as emergency conversations and adolescent scenarios. In actual deployment, a more appropriate approach is to establish custom weights based on the target population: a companion app for adolescents clearly cannot substitute a specialized evaluation with an overall score dominated by adult samples.

Aggregators such as OpenAI Hub can reduce the cost of cross-model testing and switching, but API compatibility does not mean safety-capability compatibility. The same prompt may produce different questioning styles, crisis-escalation tendencies, and refusal boundaries across models. Mental health scenarios are particularly unsuitable for treating model replacement as a purely cost-optimization exercise.

Evaluation Is Finally Moving from “Don’t Say the Wrong Thing” to “What Should You Do?”

The direction taken by MentalHealthBench is sound. It is no longer satisfied with checking whether a model has produced obviously dangerous content. Instead, it has begun measuring whether a model can understand context, respect autonomy, ask appropriate questions, and provide actionable real-world support.

But the top score of 57.3% also reminds the industry that current large models are, at most, an entry point in the mental health support chain—not therapists. They can help users organize their emotions, encourage them to seek help, provide basic information, and serve as temporary conversation partners late at night. When diagnosis, treatment, or crisis intervention is involved, however, users must still be brought back into a professional network in the real world.

More noteworthy is OpenAI’s decision to make the benchmark public, allowing external researchers to inspect the methodology and conduct independent evaluations. This will force model developers to turn “more empathetic” from a launch-event adjective into data that can be reproduced, challenged, and compared.

What will ultimately be convincing is not which model scores two or three points higher under the same judge, but whether MentalHealthBench can support cross-institutional verification: repeated testing with different scoring models, real clinicians, and region-specific crisis resources. If the rankings remain stable, the leaderboard may gradually evolve from a vendor self-test into industry infrastructure.

Sources

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: