Investigation · Original testing
AI beat the doctors. So we regraded the doctors.
In July, OpenAI’s GPT‑5.6 was reported everywhere to have outperformed physicians on health evaluations — 60.5 against the doctors’ 43.7. That comparison appears in no OpenAI publication. The 43.7 comes from a different paper about a different model, and it is not a fixed quantity: it is a number produced by one grader. We downloaded the physicians’ answers OpenAI released, and had four other frontier models — including GPT‑5.6 Sol itself — score them again.
Every independent grader we ran put the physicians higher than OpenAI’s published figure of 43.7 — between 47.1 and 51.9 on the benchmark’s own scale, using the benchmark’s own rubrics and its own length adjustment. On those numbers the headline gap between the AI system and the doctors roughly halves, from 15.3 points to somewhere between 7 and 12. The doctors’ score, in other words, is not a property of the doctors. It is substantially a property of who is holding the pen.
That is the finding. The rest of this piece is how a claim that no one published came to be reported by dozens of outlets, and what OpenAI’s benchmark is actually measuring when it says a model beat a physician.
The comparison in the headlines appears in no OpenAI document
The claim reached most readers through aggregators. A Crypto Briefing article of 11 July, syndicated verbatim by KuCoin, reported that GPT‑5.6 Sol scored 60.5 on HealthBench Professional, GPT‑5.5 scored 59.0, and physician-written responses scored 43.7 — and that “doctors reviewed medical responses, rated them blind, and the AI came out ahead. Not slightly ahead. Measurably, consistently ahead.”
Those numbers come from five separate OpenAI communications, and the story fuses them.
- The 43.7 is from the HealthBench Professional paper of 30 April 2026. It is the score for physician-written responses, measured against GPT‑5.4 in ChatGPT for Clinicians, which scored 59.0. GPT‑5.6 does not appear in that paper. Neither does GPT‑5.5.
- The 59.0 attributed above to GPT‑5.5 is therefore wrong. It is the April score for the GPT‑5.4 product harness. GPT‑5.5 actually scored 51.8.
- The 60.5 is real and current: it is GPT‑5.6 Sol’s length-adjusted score in the GPT‑5.6 preview system card. That document reports Sol 60.5, Terra 57.7, Luna 55.7 and GPT‑5.5 51.8 — and contains no physician comparison at all. No physician row, no physician baseline, no claim to have beaten one.
- The “260 physicians across 60 countries” who “reviewed over 700,000 model responses” describe OpenAI’s standing advisory programme, from a June post about GPT‑5.5 Instant. They are not the cohort behind the 43.7; the HealthBench Professional paper used 190 physicians across 50 countries.
- The “blind” comparison is real, but it is not the benchmark and it is not published. It comes from a thread by OpenAI’s Karan Singhal on 9 July, amplified by Sam Altman: physicians wrote answers with unlimited time and web access, other physicians compared them side-by-side blinded to source across five axes, and OpenAI reported the fraction rated perfectly on every axis across 20,000 axis ratings. “GPT‑5.6 Sol appeared strongest, although all GPT‑5.6 models performed significantly better than physicians.”
So the specific comparison in circulation — 60.5 against 43.7 — is a subtraction performed by the press across two documents written ten weeks apart about different models. And the strongest-sounding claim of the whole cycle, that all GPT‑5.6 models beat physicians, rests on a social media post with no paper, no methodology, and no numbers anyone outside OpenAI can check.
It is worth being precise about who did what. OpenAI’s own posts were more careful than the coverage they produced: Singhal and Altman both noted that the tasks were selected to be difficult for OpenAI’s models and that the physician baseline was a solo writing exercise. Downstream they were almost entirely dropped: across 55 articles we coded, six mentioned the adversarial selection and six the nature of the physician baseline. By 12 July the claim had reached Marc Andreessen, who wrote that “AI is already a better doctor than 99.99% of human doctors” — carried by Benzinga and syndicated to Yahoo Finance. That is the distance travelled in seventy-three days: from a rubric score on an adversarially selected test set, to a superlative about essentially every physician alive.
What the benchmark actually scores is checklist coverage
HealthBench Professional is a real piece of work, and unusually transparent by vendor-benchmark standards. It contains 525 conversations physicians had with ChatGPT for Clinicians, each graded against criteria written and adjudicated by three or more physicians. OpenAI released the full dataset — including the physicians’ own answers, which is the only reason the test below was possible.
But it is worth knowing what the score is. A system’s answer is marked against a short checklist — a median of two criteria per example, and a fifth of the set has exactly one, making the score for those examples strictly binary. The score is the fraction of available points captured. It is not accuracy, not an error rate, and the paper says plainly that it is not a real-world performance rate: the examples were deliberately enriched about 3.5× for cases where OpenAI’s own models had failed, and the authors write that “a moderate aggregate score (e.g., 45%) on HealthBench Professional can coexist with high real-world performance in typical usage.”
The physician baseline is a specific thing too. For each example, one specialty-matched physician was asked to write the best possible next chat message, with unlimited time and web access but no AI. That is a solo response-writing exercise scored on checklist coverage — not a diagnosis, not an episode of care, not a patient outcome. Their answers are also much shorter than the models’: a median of about 1,700 characters against 3,200–3,800 for the models. The benchmark’s length adjustment actually works in the physicians’ favour here, by roughly three to five points.
One detail rarely survives into coverage. On the slice the paper itself labels hard-but-realistic — good-faith questions where the model struggled — the physicians were statistically tied with every model tested, and numerically ahead of OpenAI’s flagship system. The headline margin comes overwhelmingly from the routine slice and the adversarial red-team slice. In a review of 55 articles covering these claims, that fact appeared in none of them.
We gave the physicians’ answers to four other graders
Here is the part that matters. The benchmark is graded by a model — GPT‑5.4 at low reasoning effort, an OpenAI model marking a contest that an OpenAI system wins. The obvious question is whether the physicians’ 43.7 is a fact about the physicians’ answers or a fact about that grader.
So we took the 525 physician responses OpenAI published, kept the benchmark’s own rubrics, its own scoring rule and its own length adjustment, and changed exactly one thing: the grader. Four independent frontier models scored the same text — two from Anthropic, and two from OpenAI’s own family, including GPT‑5.6 Sol, the model at the centre of the claim.
Every grader placed the physicians above 43.7. Fable 5 scored them 51.9 (95% CI 47.1–56.4), Opus 4.8 50.7 (46.1–55.1), GPT‑5.5 49.1 (44.4–53.6), and GPT‑5.6 Sol 47.1 (42.3–51.8). Three of those four intervals exclude OpenAI’s published figure. Sol’s does not — on the evidence of the newest OpenAI model alone, 43.7 cannot be ruled out, and we say so rather than reporting only the three results that make the cleaner story.
The graders agreed with each other on about 87–94% of individual examples, so this is not noise. It is a systematic difference in strictness, and it runs in a suggestive order: the two Anthropic models are most generous, then GPT‑5.5, then GPT‑5.6 Sol, then OpenAI’s published GPT‑5.4 figure — strictness rising as you move toward OpenAI’s own stack. We cannot tell you why. Shared training lineage, genuine differences in judgement, and the interaction between each model and our grading prompt cannot be separated with what OpenAI has released, and we are not going to assert a cause we cannot demonstrate.
What this does and does not establish is worth stating exactly. It does not prove OpenAI’s number is wrong; a grader has to be chosen, and theirs is defensible. It does not produce a corrected gap, because OpenAI released the physicians’ answers but not the models’, so only one side of the comparison can be re-scored — if our graders are more lenient in general, the model scores would rise too. What it establishes is narrower and still substantial: the anchor of the “beats physicians” claim moves by three to eight points depending on who grades, and not one published version of that claim carries the uncertainty.
A cheap model with a fixed template captures most of the physicians’ score on routine questions
If the score is checklist coverage, the natural question is how much of it a model can capture without doing any medicine. So we ran a second test. Haiku 4.5 — a small, cheap model — answered all 525 examples seeing only the conversation. It never saw the rubric or the physicians’ answers. It was given a fixed template — answer directly, cover the standard bases, name red flags, note that guidance varies by age and setting, ask a clarifying question, flag any doubtful premise, give next steps — and told explicitly not to attempt case-specific reasoning.
On the routine slice — 49% of the benchmark, the part most like everyday clinician use — that rubric-blind template captured 85–88% of what specialty-matched physicians with unlimited time achieved. Across the whole set it beat the physician outright on about a third of examples. Both of the graders we ran it past agree on this.
But it falls away as the cases get harder: to 38% of physician level on the hard-realistic slice and 57% on the adversarial one under the first grader, and 30% and 41% under the second. That cuts the other way, and it matters. Those slices are not measuring formatting. They demand case-specific judgement that generic thoroughness cannot fake, and the physicians hold a large lead there. HealthBench Professional’s hard slices are carrying real signal. Its routine half is substantially a test of conscientious coverage — which is exactly what a frontier chat model is trained to produce.
We also tested the most attractive explanation we had and had to throw it away. We suspected the template’s reflexive hedging would let it dodge the red-team traps that catch a committed expert. It does not: on the examples that carry a penalty criterion, the physicians and the template end up net-negative at nearly identical rates — 15% against 16% — and across those examples the physicians score almost twice as well (35.5 against 19.4). The template’s wins come from coverage on ordinary questions, not from evasion.
Then look at what a physician has to write to score well
One case is worth more than the aggregate. A clinician asked for a template for reporting an exercise stress test. The benchmark scores that answer against a single criterion worth eight points, which requires all sixteen of a list of elements — identifying information, indication, active medications, chronotropic medication status, exercise duration, METs, reason for stopping, ECG at baseline, at peak and in recovery, resolution of abnormalities and of symptoms, arrhythmias, internal and external workload, and an impression.
The cardiologist wrote a clean, usable report template of the kind that gets used in clinic. It covers about ten of the sixteen. Because the criterion is all-or-nothing, it scored 3.3. The generic template, which produced a long checkbox list covering everything and committing to nothing, scored 93.4. All four of our graders scored that physician answer identically, so this is not a grading artefact — it is what the metric is built to reward.
We have put eight of these cases online so you can judge them yourself, blinded: the question, both answers, then the rubric and every grader’s score. They are chosen to cut both ways, and include cases where the benchmark works exactly as intended — a physician delivering a genuine Amharic translation the template cannot fake, an intraocular lens power calculation with a right answer, a red-team trap the template walks straight into.
What to take away
Nothing here shows that physicians outperform GPT‑5.6, and nothing here shows OpenAI cheated. The benchmark is carefully built, the paper is candid about its own limits, and OpenAI released the data that made this critique possible — which is more than most vendors do.
What we can say is narrower and worth holding onto. The comparison that travelled the world was assembled by the press from documents that never made it. The number anchoring it moves by three to eight points depending on which model grades, and every published version of the claim states it as a fixed quantity. Roughly half the benchmark can be substantially satisfied by a cheap model with a fixed template and no case-specific reasoning. And on the hardest realistic cases, the physicians were never actually beaten.
“Beats physicians” is doing a great deal of work in that sentence. It means: produced chat responses that covered more of a short physician-written checklist, on a test set built from the questions clinicians typed into a chatbot and weighted toward the ones the chatbot got wrong, as marked by a language model. That is a real and useful thing to measure. It is not the practice of medicine, and the distance between the two is where the entire story lives.
Methods & reproducibility
We downloaded the released HealthBench Professional dataset (525 examples, MIT licence) and reimplemented the paper’s scoring rule from §4.1: each rubric criterion is marked met or not met; the raw score is captured positive points over available positive points, with negative criteria subtracting from the numerator only; the length-adjusted score is raw − 2.94×10⁻⁵ × (characters − 2000); the mean is clipped to [0,1] and reported ×100. Aggregates are reweighted to the benchmark’s true slice composition (256 routine / 78 hard / 191 adversarial) so that differences in grader coverage cannot bias them, and confidence intervals are 4,000-sample bootstraps.
Coverage was 512–525 of 525 examples per grader, the shortfall being examples a grader returned with an unmatchable id, or whose rubric carries no positive points and so has no defined score. The two OpenAI-family graders were run through the Codex CLI; the two Anthropic graders through their own API. The generic-template arm was generated by Haiku 4.5 from the conversation alone and scored by two graders under the identical protocol.
Three limitations bind the result. Grader model and grading prompt cannot be fully separated, because OpenAI’s evaluation implementation is unreleased. Only the physician arm could be re-scored, because the model responses were never released. And we are language models auditing a language-model judge — the circularity we are pointing at applies to our instrument too, which is why every grader’s agreement statistics are published alongside its score rather than only the headline number.
The single thing that would settle it is a human anchor: practising physicians grading a sample of these responses blind. We did not have one, and no published version of this claim — OpenAI’s included — rests on one either.
Disclosures & provenance
- Published
- 23 Jul 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Rating
- Overstated (2/5), rubric v1.0, rated 23 Jul 2026. Full rating card →
- Corrections
- None to date. Corrections log → · Challenge this analysis