Investigation · Benchmarks
“Medical superintelligence”: what Microsoft’s 85.5% actually beat
Microsoft says its diagnostic orchestrator solves NEJM case challenges at more than four times the rate of experienced physicians. The 85.5% is real — on Microsoft’s own benchmark. What it beat was 21 generalists working rare-disease puzzles with every reference tool of modern practice taken away, scored on a different slice of the cases, graded by the model family being tested.
On June 30, 2025, Microsoft AI published a research announcement titled “The Path to Medical Superintelligence.” Its centerpiece: the Microsoft AI Diagnostic Orchestrator (MAI-DxO) “correctly diagnoses up to 85% of NEJM case proceedings, a rate more than four times higher than a group of experienced physicians.” Mustafa Suleyman, Microsoft AI’s chief executive, told Wired the company had taken “a genuine step toward medical superintelligence,” and told The Guardian such systems are on a path to getting “almost error-free in the next 5-10 years.” The 85.5% is a real number, produced by genuinely interesting engineering. It is also a number about a very specific, very unusual test — and the sentence that carried it into the world removed most of what makes it specific. This is the figure that will be quoted when health systems, insurers and patients decide what an AI’s diagnostic judgment is worth. It deserves to be read at the resolution it was produced.
The claim, and where it traveled
The paper behind the announcement — “Sequential Diagnosis with Language Models” (arXiv:2506.22405), all fifteen authors Microsoft AI employees — states the comparison in its abstract: paired with OpenAI’s o3, MAI-DxO “achieves 80% diagnostic accuracy—four times higher than the 20% average of generalist physicians,” rising to 85.5% “when configured for maximum accuracy.” Microsoft’s blog post adds one load-bearing phrase the paper is more careful about: it says the physicians scored their 20% “on the same tasks.” Hold that phrase. It is not quite true, and the gap it papers over is most of this investigation.
The claim traveled fast, and its qualifiers did not travel with it. The Guardian headlined an AI “better than doctors at diagnosing complex health conditions”; Wired, an AI that “diagnosed patients 4 times more accurately than human doctors”; Fortune, Business Insider, the Financial Times and the Economic Times ran the same arithmetic. Most of those six accounts did carry one condition — that the physicians worked without their usual tools. None stated the other two plainly: that no patients were involved anywhere in the study, and that the physicians were scored on a different, smaller slice of the benchmark than the AI was. The skepticism arrived later and further down the readership curve, from the medical press: STAT reported that experts found the “superintelligence” framing misplaced, and the medical news service Healio quoted City of Hope’s chief AI and analytics officer calling the four-times framing “exaggerated.”
What the benchmark actually is
Credit first: SDBench is a more serious test than the multiple-choice exams that produced earlier “AI passes the medical boards” headlines. Microsoft converted 304 clinicopathological conference (CPC) cases from the New England Journal of Medicine — published 2017 through 2025 — into stepwise encounters. The diagnostician starts from a short case abstract and works turn by turn — asking questions and ordering tests, with agents allowed to batch many in a turn (MAI-DxO limited itself to five questions and three tests) — against a “gatekeeper” model (OpenAI’s o4-mini) that holds the full case file and reveals findings only when explicitly requested; for questions the file doesn’t cover, the gatekeeper invents “realistic synthetic findings,” with no indication they are synthetic. A final diagnosis is scored by a “judge” — OpenAI’s o3, applying a physician-authored rubric, with a score of at least 4 on a 5-point scale counted as correct.
Two properties of that design matter for everything downstream. First, NEJM CPC cases are teaching puzzles, selected precisely because they are hard and unusual — the paper’s own limitations section says the case distribution “does not match that of a real-world deployment scenario,” that no patient in the set is healthy or has a benign, self-limiting condition, and that the authors therefore “could not measure false positive rates.” A test with no healthy patients cannot tell you how often a diagnostician cries wolf — which, in a screening-and-referral world, is half of what diagnostic skill is. Second, correctness is a language model’s judgment. Microsoft validated the o3 judge against its own in-house physicians on 56 cases and found substantial agreement (Cohen’s κ of 0.70 on the AI’s diagnoses, 0.87 on the humans’) — and, to be fair in both directions, where they disagreed, the physicians mostly found the judge too strict. But an OpenAI model grading answers produced by OpenAI models, inside a study written entirely by one company about its own product, is a closed loop from claim to verdict.
Not the same test: 304 cases for the AI, 56 for the humans
Here is the condition no headline carried. The paper splits its 304 cases into a 248-case “validation” set, used while the team developed and tuned MAI-DxO, and a 56-case held-out test set — the most recent cases, from 2024–2025 — kept hidden until the methods were final. The 21 physicians were only ever run on the 56 held-out cases. The AI’s headline numbers — 78.6% for off-the-shelf o3, 79.9% and 85.5% for MAI-DxO configurations — are, in the paper’s own words, “evaluated on all 304 NEJM cases (including the 56 test set cases), while physician performance is shown only for the held-out 56 test set cases.” The “four times higher” is a ratio between two numbers measured on two different case sets, 248 of which the AI side had used during development. Microsoft’s blog rendered this as physicians scoring 20% “on the same tasks.”
The steel-man deserves stating in full. On the 56 held-out cases the AI systems did not get worse — if anything they “tend to perform better on this subset” — and most of those recent cases postdate the models’ training cutoffs, which makes memorization an unlikely explanation and is a genuinely better contamination control than most benchmarks attempt. Better still, a like-for-like number does exist: an appendix ablation reports MAI-DxO with o3 at 83.9% on the same 56 cases the physicians took — which, against their 19.9%, is a same-set gap of about 4.2 times. The honest comparison was available, and it is, if anything, slightly stronger than the printed one. That is what makes the presentation hard to excuse rather than easy: the abstract and the blog ran the cross-set arithmetic anyway — a 304-case number set against a 56-case number — and the blog captioned it “the same tasks.” The gap’s direction is real. The sentence describing it is still wrong about what was compared.
The physicians were playing a different game
The 19.9% physician average is the study’s most quoted and least examined number. It was produced by 21 physicians — 17 in primary care, 4 hospital generalists, median 12 years in practice (the blog says “5-20 years”; the paper says the interquartile range runs 6 to 24) — each completing on average 36 of the 56 cases through the same chat interface as the models. They were, per the paper, “explicitly instructed not to use external resources, including search engines… language models… or other online sources of medical information.” The authors had a defensible reason: the NEJM cases are findable online, and an unrestricted physician could simply look the answers up. But the same limitations section concedes what the restriction costs: in reality physicians “are free to use such tools,” along with electronic records, guidelines, colleagues and textbooks — and generalists facing a CPC-grade mystery would normally refer it to a specialist, not solve it solo in the 11.8 minutes per case the paper reports. The study removed every scaffold of real practice from the humans while benchmarking the machine at full strength; Microsoft’s blog describes this as enabling “a fair comparison to raw human performance.” Raw, yes. Representative of “doctors,” no — a point made bluntly by clinicians within days: MIT’s David Sontag noted in Wired that the tool restriction “may not be a reflection of how they operate in real life,” and City of Hope’s Nasim Eftekhari, more directly: “I feel that is exaggerated.”
Off-the-shelf o3 already scores 78.6% on this benchmark. Most of what beat the physicians was not Microsoft’s orchestrator. It was the restrictions — and the base model.
Two more numbers put the 19.9% in proportion. The best individual physician scored 41% — roughly double the mean, still under the restrictions — which is a wide spread for a baseline presented as what “experienced physicians” achieve. And plain o3, with no orchestration at all, scored 78.6%: the celebrated “system” adds between one and seven points to a foundation model Microsoft did not build, depending on configuration. That is honest, interesting engineering — MAI-DxO’s real contribution is doing it at much lower simulated cost — but it means the headline is mostly a claim about o3 plus the test conditions, not about a Microsoft diagnostic breakthrough.
“More cost-effectively than physicians” — for a different configuration
The announcement’s second claim is economic: MAI-DxO “delivered both higher diagnostic accuracy and lower overall testing costs than physicians or any individual foundation model tested.” The costs are synthetic — each visit is priced at a flat $300 and each ordered test is mapped to a CPT code and priced from a single US health system’s 2023 price table; the authors call the estimates “first-order approximations” — but the more basic problem is that the accuracy claim and the cost claim belong to different machines. The configuration that undercuts the physicians ($2,396 per case against their $2,963) is the budgeted one, at 79.9% accuracy. The 85.5% configuration — the one in the headline — runs multiple panels in parallel and costs $7,184 per case, about 2.4 times the physician average. There is no configuration in the paper that is simultaneously the most accurate and cheaper than the physicians.
A year later: the claim escalated; the evidence didn’t move
Investigations age, so here is what twelve months did. The paper, which Microsoft said in June 2025 was “in the process of submitting… for external peer review,” has no peer-reviewed version we could find as of July 2026 — the arXiv record stops at v2, July 2, 2025. SDBench itself was never released, so we could find no independent reproduction of the 85.5%, and no clinical trial of MAI-DxO appears on ClinicalTrials.gov in a registry search on July 27, 2026. What did move was the rhetoric. In November 2025, announcing an in-house “superintelligence team,” Suleyman wrote that MAI-DxO “managed to reach 85% across the Case Challenges” while “human doctors max out at about 20%” — a rendering the paper contradicts on its face, since the paper’s own best doctor scored 41% and its authors called the physician baseline “a first-order approximation.” By March 2026, launching the consumer product Copilot Health, Suleyman wrote that we are “approaching the dawn of medical superintelligence”; the product’s own disclaimer states it is “not intended to diagnose, treat, or prevent diseases or other conditions.” Meanwhile a September 2025 paper co-authored by two of the SDBench authors (with Eric Topol) stress-tested six medical benchmarks and concluded that static benchmark scores “provide little insight into how models behave under real-world uncertainty” — as close as the field comes to an in-house second opinion on what an NEJM-case score can carry.
What would move this rating
Every step is concrete and available to Microsoft. Publish the peer-reviewed version. Release SDBench so the 85.5% can be independently reproduced. Promote the like-for-like comparison from the appendix to the headline — every configuration scored on the same 56 cases the physicians took — and run a physician arm equipped the way physicians actually work, with references, colleagues and their own AI tools, which is the comparison a deployment decision needs (several clinicians suggested the even better one: physicians with MAI-DxO versus physicians without). Above all, test the system prospectively on real, undifferentiated patients — everyday case mix, healthy people included, false positives counted. The moment evidence of that shape exists, this rating gets re-scored in public, in either direction. The pattern to watch is the one this site exists to track: a claim that grows in public while its evidence stands still. For the reader’s own toolkit, this claim is a textbook case for the seven questions we ask of any “AI beats doctors” study, and its row sits alongside fourteen others in our scoreboard of what these studies actually measured.
The bottom line
Believe the 85.5% — as a score on Microsoft’s own simulated, curated, machine-graded benchmark, where plain o3 already scores 78.6%. Do not believe, on this evidence, that an AI diagnoses “four times better than doctors”: the physicians took a harder version of a different test, without the tools that make them doctors, and nobody — not a patient, not a regulator, not an independent lab — has ever been able to check the number since. The interesting engineering is real. So is the gap between what was measured and what was said. When the peer-reviewed paper, the released benchmark, or the prospective trial arrives, we’ll read it here — and update this rating in public if it earns a change.
Primary sources: Nori et al. (Microsoft AI), “Sequential Diagnosis with Language Models” (arXiv:2506.22405, 2025) · Microsoft AI, “The Path to Medical Superintelligence” (2025) · Microsoft AI, “Towards Humanist Superintelligence” (2025).
Claim as circulated: The Guardian · Wired · Fortune · Business Insider.
Critical coverage: STAT News · Healio.
Ground Truth is independent of, and unaffiliated with, Microsoft and OpenAI. Corrections are made in public — see our policy.
Disclosures & provenance
- Published
- 10 Jul 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Rating
- Overstated (2/5), rubric v1.0, rated 10 Jul 2026. Full rating card →
- Corrections
- None to date. Corrections log → · Challenge this analysis