What changed in the format

Most medical benchmarks before this month asked a model to pick A, B, C or D. HealthBench, released by OpenAI on May 13, asks a model to answer a multi-turn conversation and then scores the reply against a rubric written by a physician for that specific conversation. There are 5,000 conversations, 48,562 unique rubric criteria, and 262 physicians behind them. The median conversation has eleven criteria, and the range runs from two to 48.

Each criterion carries a nonzero point value between minus 10 and 10. Positive criteria describe things a good answer should do, such as asking a follow-up question about knee pain to narrow the diagnosis. Negative criteria describe things it should not do. A model grader reads each criterion independently, decides whether the response meets it, and awards full points or none. The example score is the sum divided by the maximum available, clipped to the range from zero to one, so a response can be scored negative before clipping if it trips enough penalties.

That is the whole grading mechanism. There is no free-text judgment of overall quality and no pairwise preference. The grader answers a long series of yes or no questions that a physician wrote in advance. The paper cites earlier rubric evaluations it builds on, so the idea did not originate here. Still, 48,562 hand-written criteria on realistic health conversations is a much larger commitment to the format than we have seen before.

How the grader is checked against physicians

The obvious worry is that a model grader with a physician's rubric is still a model grader. The authors address this with a meta-evaluation on a subset called HealthBench Consensus. Consensus criteria are 34 pre-written criteria, such as whether a response gives a clear emergency referral when one is warranted. A consensus criterion is attached to a conversation only when a majority of reviewing physicians, at least two, agree it applies. The 34 criteria appear 8,053 times across the dataset.

For those criteria, physicians also graded actual model responses as met or not met. That produced over 60,896 meta-examples, an average of 1,791 per criterion. The authors then compare the model grader, GPT-4.1, against physicians using macro F1, which averages the F1 for the met class and the not-met class so that rare failures count as much as common successes.

The result is a table by theme. Physician-to-physician macro F1 ranges from 0.569 on response depth to 0.730 on health data tasks. The model grader lands at 0.572 and 0.683 on those same themes. Across the seven themes, the grader sits between the 37.5th and the 88.2nd percentile of individual physician scores. The honest summary is that the grader agrees with physicians about as well as physicians agree with one another, and physicians do not agree with one another all that well. A macro F1 in the 0.6 range means real disagreement on whether a criterion was met, whichever grader you use.

This validation covers only the 34 consensus criteria, which are 14 percent of criterion instances. The remaining 86 percent are example-specific criteria that were written once by one physician and graded by the model without a second human check. The paper is open about this. HealthBench Consensus is described as having greater precision and lower recall for finding failures, and the full benchmark as the reverse.

What the scores say

GPT-3.5 Turbo scores 16 percent, GPT-4o 32 percent and o3 60 percent. On HealthBench Hard, a subset of 1,000 examples chosen for difficulty, the top score is 32 percent. GPT-4.1 nano scores higher than GPT-4o at 25 times lower cost. Of the non-OpenAI models, Grok 3 and Gemini 2.5 Pro were reported as markedly ahead of Claude 3.7 Sonnet and Llama 4 Maverick. Almost 40 percent of all rubric items relate to completeness, which is the axis on which models most often lose points.

The physician baseline is the part we keep rereading. Physicians writing answers from scratch, without model help, scored relatively weakly. When given reference responses from September 2024 models, physicians improved on them more often than they made them worse, 56.2 percent against 39.8 percent. Given references from April 2025 models, physicians were as likely to worsen a response as improve it, 46.8 against 47.7 percent. The authors flag that unassisted physician responses were much shorter and that scores correlate with length, so we would resist reading this as models outperforming doctors at medicine. The narrower claim is that on this rubric and this task framing, a long careful model answer is hard for one physician to improve by hand.

What it costs to build one of these

The paper gives no dollar figure. It gives enough to estimate the effort. The physician cohort worked over eleven months. 1,021 physicians initially expressed interest and were screened down to 262 who passed quality checks, with 26 specialties, 60 countries of practice experience and 49 languages between them. A physician lead and a group of physician advisors wrote training material, ran live trainings, reviewed tasks and gave individual feedback. Each of the 5,000 conversations then needed a rubric written by hand, and the consensus criteria needed multiple independent physician judgments per example.

So the cost of a rubric benchmark is mostly the cost of expert time, and it scales with the number of examples rather than with model queries. Multiple-choice benchmarks front-load a smaller amount of expert effort and then run for free. A rubric benchmark front-loads a lot more and then pays a model grading cost on every evaluation, with the meta-evaluation as a separate line item. That is why we expect the format to spread in domains where experts can be recruited at scale and to stay rare elsewhere.

What we would want next

Two things. First, a meta-evaluation of the example-specific criteria, not only the consensus ones. Even a few hundred physician-graded meta-examples drawn at random from the 48,562 would tell us whether the 0.6 to 0.7 macro F1 holds for criteria that were written once and never reviewed. Second, an experiment on grader choice. The paper reports GPT-4.1 as the grader. If a different lab's model grades the same responses and the leaderboard order changes, that is a fact the field needs before rubric scores get quoted as if they were exam results.

We think rubric grading will become the default for open-ended evaluation, because it is the only format we know that produces a score, an explanation of the score, and a per-criterion audit trail at the same time. The HealthBench paper is a good template for how to report the validation. It is also a reminder that the validation is the expensive part, and the part most likely to get skipped when the format is copied.

Sources

  1. Arora et al., HealthBench: Evaluating Large Language Models Towards Improved Human Health (arXiv 2505.08775)