MentalHealthBench evaluates AI across realistic mental health conversations

An open benchmark created with more than 80 licensed experts assesses whether AI responses to mental-health conversations are safe, contextual and useful.

OpenAI visual for MentalHealthBench
Image: OpenAI

MentalHealthBench is an open benchmark developed with more than 80 licensed psychologists and psychiatrists from 22 countries, representing 19 languages and nearly 20 mental-health specialties. It evaluates how AI systems respond to synthetic but realistic conversations ranging from everyday stress to high-acuity situations and emergencies.

Each conversation was reviewed by at least three experts, who defined criteria covering safety, context, user agency and actionable guidance. OpenAI is releasing the benchmark so researchers can examine and extend the methods. The work measures model behaviour; it does not make ChatGPT a substitute for therapy, diagnosis or professional care.

About this Item

Geography
Language