OpenAI Launches HealthBench to Evaluate AI in Healthcare
Home > Health > Article

OpenAI Launches HealthBench to Evaluate AI in Healthcare

Photo by:   Unsplash
Share it!
Sofía Garduño By Sofía Garduño | Journalist & Industry Analyst - Tue, 05/13/2025 - 10:53
DIA assistant

OpenAI has introduced HealthBench, a new benchmark aimed at evaluating the performance of AI systems in healthcare settings through realistic, physician-reviewed scenarios. Developed with input from 262 doctors across 60 countries, HealthBench includes 5,000 health-related conversations and a comprehensive rubric for scoring model responses.

“We hope that the release of this work guides AI progress towards improved human health,” says Karan Singhal, Health AI Team Lead, OpenAI.

The benchmark is part of OpenAI’s broader initiative to ensure that evaluations of AI systems in health reflect real-world use, physician judgment, and ongoing room for improvement. The conversations, designed to simulate interactions between users and large language models (LLMs), cover a range of medical specialties and user types. They are multi-turn, multilingual, and were produced using both synthetic generation and adversarial human testing to better replicate actual clinical communication.

Each model response in HealthBench is graded against physician-authored criteria specific to the case. These criteria are weighted by importance and span 48,562 individual benchmarks, evaluating aspects such as clinical accuracy, communication quality, and the ability to seek context. The final score reflects how many criteria a model met compared to the total available.

To provide further depth, OpenAI released two additional sets within the benchmark: HealthBench Consensus and HealthBench Hard. The Consensus subset includes 3,671 examples validated through physician agreement, offering near-zero error baselines. HealthBench Hard features 1,000 challenging cases where AI systems underperform, designed to promote further advances in model development.

Performance on HealthBench was evaluated using OpenAI’s latest GPT-4.1 model, which also served as an automated grader. To validate grading reliability, OpenAI conducted meta-evaluations, comparing model grading to physician assessments. The results showed that agreement rates between the model and physicians were on par with inter-physician agreement, supporting the use of automated tools for scalable benchmarking.

The benchmark also introduces a framework for analyzing model behavior across seven thematic categories, including emergency care, uncertainty handling, and global health. Each theme provides context-specific criteria to capture a range of clinical challenges.

HealthBench is part of OpenAI’s effort to create rigorous and transparent tools for evaluating AI in high-impact domains. According to the organization, while AI models have shown significant progress and can even outperform human experts in some cases, there is still a need for improvement in key areas, such as understanding underspecified queries and ensuring reliability in critical scenarios.

OpenAI has made the entire HealthBench dataset and evaluation tools publicly available through GitHub, encouraging feedback and collaboration from the broader research and healthcare communities. The goal is to support the development of AI systems that contribute meaningfully to global health outcomes.

“Where we go from here is on us. Point any model you want at HealthBench, publish the gaps, patch them, rerun,” says Jonathan Slotkin, Chief Medical Officer, Contigo Health.

Photo by:   Unsplash

You May Like

Most popular

Newsletter