Investigation · Benchmarks
OpenEvidence’s “perfect 100% on MedQA” covers 660 of the test’s 1,273 questions
On 3 September, OpenEvidence announced that its Darwin model is “the first AI in history to score a perfect 100% on MedQA, the leading independent benchmark of medical AI.” Darwin was scored on 660 of the 1,273 questions in MedQA’s test set. The 613 left out include every question on which two or more physicians in Google’s 2024 review of the test, answering before they saw the official answer, chose a different one. The last 18 were removed in an August review that re-examined only the questions some model had got wrong. On the same 660, two other AI models we tested score 2.4 and 9.6 percentage points higher than on the full test. OpenEvidence’s methods post, linked from its announcement page, describes the 660 and the review; the press release does not mention them. Darwin’s released answers to all 660 scored questions match MedQA’s original answer key, and it is a strong system on other measures. Its perfect score belongs to a selected half of the test.
MedQA is a set of multiple-choice questions in the style of the US medical licensing exam. Its test split, the part models are scored on, has 1,273 questions, each with four options in the standard version used here (Jin et al.), and each with one official answer, the option MedQA’s answer key marks correct. On 3 September 2026 OpenEvidence, which makes an AI search tool for clinicians, introduced a research-preview model called Darwin, and its press release called Darwin “the first AI in history to score a perfect 100% on MedQA, the leading independent benchmark of medical AI.”
That sentence is the claim we rate, and we rate it Overstated (2/5). Darwin’s answers to every question it was scored on match the official answers; we checked all 660 of its released answers against them. But 660 questions are 51.8% of the test. The selection kept the questions on which Google’s physician reviewers most often matched the official answer before seeing it, and its final step was a review of only the questions some model had missed. On OpenEvidence’s own figures, the perfect score is two questions ahead of the next model.
OpenEvidence published the method behind the number, which is why this analysis was possible: a methods post that calls the test “a medically reviewed version of MedQA,” its reviewers’ annotations, and Darwin’s answers to all 660 questions. The qualifier stayed in the methods post. The press release, the announcement page and the X post say “100% on MedQA,” and the share-card image shows “MedQA … 100.0%.”
Darwin’s answers to the questions left out are not public, so we measured what the selection does to two models anyone can run, scoring each against the official answers. Google’s Gemini 3.7 Flash, one of the models on OpenEvidence’s own chart, gets 97.2% of the full test right and 99.5% of OpenEvidence’s 660. The open-weights Qwen3.8 27B, a smaller model whose weights anyone can download, goes from 83.3% to 92.9%. In a second experiment we regraded Darwin’s released answers on another benchmark, HealthBench Professional, with that benchmark’s own grader. OpenEvidence’s figure for Darwin largely holds up, and Darwin’s answers score far above physicians’ written answers to the same tasks.
The claim, and where the qualifier went
OpenEvidence published the result in several places on launch day. Of the five in the table below, one carries the qualifier.
Where OpenEvidence published the MedQA result on 3 September 2026
| Where | What it says | Says the set was reviewed or reduced? |
|---|---|---|
| Methods post | “the first AI in history to score a perfect 100% on a medically reviewed version of MedQA”; in the results, “it answers all 660 questions correctly” | Yes. Describes the method and the 660. |
| Announcement page | “the first AI in history to score a perfect 100% on MedQA, the leading independent benchmark of medical AI”; in the body, “MedQA—the leading fully-independent medical AI benchmark” | No. Links to the methods post through the words “publishing the results.” |
| Press release, Business Wire | The same sentence; in the body, “the first AI model in history to achieve a perfect score on MedQA” | No. No link to the methods post, though it says OpenEvidence is “publishing the results alongside Darwin’s explanation for every answer.” |
| X post | “A perfect 100% on MedQA, the first AI in history to do it.” | No. Its follow-up post links the press release. |
| Share card (the preview image for the announcement and the methods post) | “MedQA / USMLE-style board questions / Accuracy / 100.0%” | No. |
The press mostly repeated the version without the qualifier. Of 17 articles we could read in full and check at source, 12 gave the 100% with no qualifier, 3 gave “660” without saying the questions had been selected, and 2 described a reduced set. Unite.AI reported the August review as “covering every question any evaluated model answered incorrectly,” and Superpower Daily called it “a reduced and re-annotated evaluation—not the original 1,273-question test split.”
The 17 articles, and how each reported the score
- No qualifier: Fierce Healthcare (3 Sep); Digital Health News (4 Sep); explainx.ai (4 Sep); DataNorth (4 Sep); iatroX, a competitor of OpenEvidence, in a cautionary piece headed “The denominator belongs beside the percentage” (6 Sep); Becker’s Hospital Review (8 Sep); HLTH Insights (9 Sep); Digital Health News, a second article (9 Sep); Nelson Advisors (13 Sep); Grid Health, in a piece critical of self-graded benchmarks (22 Sep); Digital Health Wire (24 Sep); Wowtale (28 Sep).
- Gave “660” without saying the set was selected: AI/TLDR (3 Sep); TechTarget (4 Sep); CompleteAI Training (5 Sep).
- Described a reduced or reviewed set: Unite.AI (3 Sep); Superpower Daily (3 Sep).
- Not counted: articles we could not read in full at source (STAT, MobiHealthNews), commentary that did not restate the score, and one item whose subject was ambiguous.
How 1,273 questions became 660
MedQA’s answer key has known errors, and the best map of them is Google’s. For its 2024 Med-Gemini paper (Saab et al.), at least three US primary care physicians answered each test question without seeing the official answer, were then shown it, could revise their answer, and could flag missing information. Google counted a label error, a sign that the official answer is wrong, when a physician’s final answer left out the official one, and ambiguity when it included more than one option. Google published all 3,822 ratings. Resampling three-physician panels from those ratings, it estimated that 7.4% of the questions are unfit for evaluation under a unanimous vote, and “up to 20.9%” under a looser majority vote. It reported its model’s unfiltered accuracy, 91.1%, before the filtered figures: 91.8%, and 92.9% under the majority vote.
OpenEvidence built its test in two steps: first from those annotations, then with its own physicians’ review. According to the methods post, the first step excluded a question “flagged for missing information or ambiguity when at least two of the three physicians agree,” and a question “flagged as a label error if any physician disagrees with the original answer.” The second is described in the methods post’s own words: “To catch residual noise, three OpenEvidence physicians conducted a second review pass in August 2026 over every question that any evaluated model … answered incorrectly on the resulting set, applying the same criteria.” The steps removed 595 and 18 questions. That left 660, and 613 out: 48.2% of the test, more than twice the largest share Google estimated as unfit, although another defensible rule could draw the line elsewhere.
Google’s data also shows which questions went. Answering before they saw the official answer, Google’s physicians chose it, and only it, in 81.8% of their ratings on the 660 questions OpenEvidence kept, and in 28.9% of their ratings on the 613 it left out.
Three more facts from the same file:
- Every question that two or more of Google’s physicians answered differently from the official answer, before seeing it, was left out. None of the 660 has two such physicians.
- All 167 questions on which any Google physician flagged important missing information were left out.
- 146 of the 613 drew no objection from Google’s physicians once they had seen the official answer: each physician’s final answer was the official one alone, and no one flagged missing information or said no option was correct. Sixteen of them were among the 18 removed in August, 13 after at least one of OpenEvidence’s reviewers flagged them. The other 130 went in the first step. On all but one of those, two or more of Google’s physicians had chosen a different answer before they saw the official one.
A question on which physicians answer differently from the official answer on a first reading may be flawed, or it may just be hard. Google’s reviewers assessed each question again after showing the official answer, which helps tell the two apart, although agreement after seeing the answer does not prove a question sound. The 130 questions above passed that second look and were left out anyway. As far as we can reconstruct it (next section), the first step selected on how physicians answered before they saw the official answer.
We could not reproduce the set from the published rule
OpenEvidence’s stated first-step rule can be applied to Google’s public annotations, and we did that. We could not reproduce its set. Google’s paper and its analysis code define a label error as a physician whose answer, after seeing the official answer, does not include it. By that definition, at least 79 of the 660 questions OpenEvidence kept have a physician who disagreed with the official answer, and the stated rule says a question is excluded “if any physician disagrees.” None of those 79 was among the questions re-examined in August. We tried 1,320 combinations of reasonable readings of the rule, varying how each flag and each physician’s answer is counted. None reproduced the 678 questions implied by OpenEvidence’s account (the final 660 plus the 18 removed in August).
A rule OpenEvidence did not describe comes close. Drop a question if two or more of Google’s physicians chose a different answer before seeing the official one, or if any of them flagged important missing information. That rule, which is our reconstruction, keeps all 660 final questions. It differs from those 678 on 10: it would drop 8 questions that were in OpenEvidence’s set, all 8 later removed in August, and keep 2 that OpenEvidence left out (11 under one reading of a question whose four physicians split 2–2). It needs only Google physicians’ answers, not Darwin’s, which fits a selection made without reference to any model. A rule’s inputs cannot show how its thresholds were chosen, and we found no sign of tuning to Darwin. It is also a different filter from the one described: it selects on whether physicians could answer a question before seeing the official answer, which reflects both how hard a question is and whether it is flawed.
OpenEvidence may have used a field or a reading we did not try. It has not published the code or the exact rule.
The August review: 28 questions re-examined, 18 removed
The second step is documented in a file OpenEvidence released, uploaded to its public storage on 24 August 2026. It has 84 rows covering 28 questions, three reviewer slots per question; one slot, for question 215, reads “N/A” throughout. According to the methods post, these are the questions that at least one of five models answered wrongly after the first step. The five are Darwin, Claude Fable 5, Claude Opus 5, GPT-5.6 Sol and Gemini 3.7 Flash. The other 650 questions in the final set were not re-examined.
Of the 28, 18 are not in the final 660. Every question any reviewer flagged for removal was removed, and none of the 10 kept had a flag. In 11 cases one physician’s flag was enough. For a disputed official answer, that is the stated rule. For missing information or ambiguity, the stated threshold is two of three.
The 28 questions OpenEvidence’s physicians re-examined in August 2026
| What the three reviewers recorded | Questions | Google question numbers | Outcome |
|---|---|---|---|
| Two of three flagged it | 4 | 21, 44, 905, 964 | Removed. Meets the stated criteria. |
| One marked the official answer as wrong | 9 | 64, 215, 226, 322, 416, 439, 831, 887, 984 | Removed. Meets the stated label-error rule, under which one physician is enough. |
| One flagged missing information or ambiguity; no one disputed the official answer | 2 | 560 (missing information), 689 (ambiguity) | Removed. Below the stated threshold of two physicians. |
| No flag from any of the three; annotations identical to the 10 kept | 3 | 290, 852, 1131 | Removed. No stated criterion applies, as we read the file. |
| No flag | 10 | 73, 541, 569, 796, 806, 853, 901, 1141, 1142, 1161 | Kept. |
answer_correct column as the physician’s judgement of the official answer.- The file has no column saying which model missed which question, so it cannot show whose errors each removal affected.
- One of the three ratings for question 215 reads “N/A” in every field.
So, as we read the released file, five of the 18 removals are not explained by the criteria OpenEvidence published: three carry the same annotations as the 10 questions kept, and two had a single flag of a kind the stated rule says needs two. The file does not say why they were removed. OpenEvidence has not published its reasons for those five removals, or Darwin’s answers to any of the 18.
Every model error triggered review, Darwin was scored in a single pass, and Darwin got all 660 kept questions right. So, by OpenEvidence’s own account of its method, Darwin’s score on the 678 questions before the review lay somewhere between 97.3% (660 of 678) and 100%. Whether Darwin missed any of the 18 decides whether its score was perfect before the review. Nothing published settles it.
The review could raise the other models’ scores as well as Darwin’s, since any model’s error could send a question to it. Its net effect on each model is unknown, because a removed question may also be one that model answered correctly. On OpenEvidence’s account, every rival error left in the final set falls on the 10 questions the reviewers kept without a flag. Reviewing only the questions models get wrong has a precedent: Vendrow and colleagues used it in 2025. They noted its limit: the errors it finds are “only a lower bound,” because a flawed question that every model answers the way the key does is never examined. For a perfect-score claim, that limit matters. Vendrow and colleagues relabelled questions with wrong keys as well as removing bad ones. OpenEvidence’s review keeps the original answer key and can only remove questions; for a model that missed a removed question, that takes an error out of its count.
OpenEvidence’s June 2026 files
The public storage folder that holds the August review file also holds files from June 2026: a PDF, stored on 22 June (UTC), headed “Generated: June 2026 / Total Questions: 667”; and a chart, in three versions, titled “Performance on MedQA,” showing “OpenEvidence 100.0,” Opus 4.8 at 99.4, GPT-5.5 at 99.1 and Gemini 3.1 Pro at 98.7, with the note “All model evaluations run on 6/21/2026.” The June set contains 659 of the final 660 questions, 7 of the 18 removed in August, and one other question. All 667 are marked correct, including the 7 later removed in August, so the OpenEvidence system that ran in June matched the official answer on those 7. The methods post mentions neither June nor 667, we have not found the June chart published anywhere, and the files do not say which OpenEvidence system produced them.
The June files help OpenEvidence on one point: before the August review, an OpenEvidence system the files do not identify had already scored 100.0 on a near-identical 667-question set, so the August review did not create the first perfect OpenEvidence score on these questions. They also complicate “first in history”: unless the June system was an early version of Darwin, an OpenEvidence system had already scored 100.0 on a near-identical set. And they show the test changing between June and August in a way the methods post does not describe. Eight questions missing from the June set, each answered differently from the official answer by two or more of Google’s physicians before they saw it, were among those OpenEvidence’s physicians re-examined in August, and all eight were then removed.
What the selection does to other models’ scores
OpenEvidence has not published Darwin’s answers to the questions it left out, so Darwin cannot be scored on the full test. The effect of the selection can be measured on models anyone can run. We chose one of the models on OpenEvidence’s own chart and a smaller open-weights model, one whose weights are public so anyone can rerun it. We ran both on all 1,273 questions and scored them against the official answers: Gemini 3.7 Flash through Google’s API at default settings, and Qwen3.8 27B run locally, with its reasoning mode off. Neither is Darwin.
On OpenEvidence’s 660, Gemini 3.7 Flash scores 99.5%, close to the 99.2% OpenEvidence reports for it. On the full test it scores 97.2%. From the exact counts (1,237 of 1,273 against 657 of 660), the selection adds 2.4 percentage points to a leading commercial model, and 9.6 to a smaller open one, which goes from 83.3% to 92.9%. For comparison, Google’s own filtering raised its model’s accuracy by 0.7 points, or 1.8 under the looser vote.
Both models score lowest on the 18 questions removed in August: Gemini 3.7 Flash answered 13 correctly and Qwen3.8 27B 7. That is partly expected, because the August review looked only at questions some model had missed, Gemini 3.7 Flash among them, and a question with a wrong official answer also produces low scores. On the 10 questions the review kept, both models scored 9 of 10.
Some of the full-test gap may be MedQA’s own fault: a model can be marked wrong because the official answer is wrong, which is why cleaning the answer key is reasonable. We have not adjudicated the models’ disagreements one by one, so we cannot size that share. The 130 questions removed in the first step although every Google physician accepted the official answer after seeing it give a cleaner comparison, though that acceptance does not prove every one of them sound. Qwen3.8 27B scores 80.0% on them, against 92.9% on the 660 kept (Fisher exact test, p < 0.001). Gemini 3.7 Flash answers all 130 correctly: each of its misses in the half left out falls on a question Google’s physicians objected to after seeing the official answer, or on one re-examined in August. So for the stronger model the selection’s lift appears to come from removing questions that Google’s or OpenEvidence’s physicians had doubts about, and for the weaker one also from removing questions whose official answer was accepted and which it finds harder. Our guide to “99% accurate” claims explains why the first question to ask of any such score is who was left out of the denominator.
In that sense the 100% is a cherry-picked number: a score on a subset on which both models we ran, and Google’s own physicians, do better than on the full test, reported under the name of the whole test. Nothing here shows that anyone chose questions to favour Darwin. The first step, by OpenEvidence’s description and by our reconstruction, used only Google physicians’ annotations, and the August review re-examined every model’s errors. The effect is familiar to benchmark builders: the authors of GPQA, whose main and diamond subsets were selected using expert validators’ agreement and non-experts’ errors, mark the validators’ accuracies on those subsets as “skewed by selection effects.”
Two questions ahead
On the final 660, OpenEvidence reports Claude Fable 5 at 99.7%, Gemini 3.7 Flash at 99.2% and GPT-5.6 Sol at 99.1%: 2, 5 and 6 wrong answers. Its results chart starts its vertical axis at 98%, so a gap of a few questions fills most of the chart’s height.
Darwin had no wrong answers on the 660; Fable 5’s reported 99.7% implies two. On those paired answers an exact McNemar test gives p = 0.50 (Fisher’s exact test gives the same): the nominal test does not establish that Darwin is better, or that the two are equivalent. A model that truly answers 99.7% of such questions correctly would score 660 of 660 about one time in seven. The exact (Clopper–Pearson) 95% interval around Darwin’s own score runs from 99.44% to 100%. These are nominal calculations. They do not account for the set having been finalised after the models’ errors were known, and without the answers from before the August review we cannot say how the review changed the gap.
The other three numbers
OpenEvidence’s methods post says Darwin is “state-of-the-art on every leading independent benchmark of medical AI.” The press release and the announcement give three more figures, each as a percentage: MedXpertQA 72.8%, HealthBench Professional 82.7% and NOHARM 87.2%. OpenEvidence released Darwin’s answers on all three, so each can be checked in part.
HealthBench Professional: the number largely holds
HealthBench Professional, from OpenAI, grades answers to real clinician queries against rubrics written by physicians. OpenEvidence’s 82.7 rests on three choices. It covers 410 of the 525 tasks: the 115 conversations with more than one exchange were dropped because, the methods post says, their earlier assistant turns are “significantly out of distribution relative to how OpenEvidence responds,” and the dataset labels 68.7% of them difficult, against 46.3% of the tasks kept. It was graded by Claude Opus 4.8; the benchmark’s paper sets GPT-5.4 at low reasoning as its default grader. And it is unadjusted for answer length, although the paper says “the primary score we report is adjusted for final response length”: longer answers have more chances to meet rubric items, so the adjustment subtracts a small penalty for each character beyond 2,000. OpenEvidence argues that its answers are long because they survey the literature, and the benchmark’s authors add a caveat in its favour: they estimated the penalty from models’ average answer lengths of no more than 4,000 characters, and note that “high-verbosity responses often exceeded the range used to estimate this length penalty.” Darwin’s answers average about 6,600 characters before their reference lists, beyond that range. OpenEvidence’s chart prints the adjusted scores too: Darwin 68.6, Claude Fable 5 67.6. On the benchmark’s primary score the lead is 1.0 point, against 12.1 unadjusted.
Because OpenEvidence released Darwin’s 410 answers, we could regrade them (our second experiment) with five graders, using the benchmark’s rubrics and scoring rule: GPT-5.4 at low reasoning, the benchmark’s own default grader; Gemini 3.1 Pro; the open-weights Qwen3.8 27B, run locally, which anyone can rerun; and GPT-5.6 Sol and GPT-5.5. We graded the physicians’ written answers to the same 410 tasks, which the benchmark publishes, fresh with the first three. GPT-5.6 Sol and GPT-5.5 had graded those same answers, with the same prompt, for a July investigation, and we reuse those grades. The analysis plan was written down before any grading.
Darwin and the physicians on the same 410 HealthBench Professional tasks
| Grader | Darwin, unadjusted | Darwin, adjusted | Physicians, unadjusted | Physicians, adjusted | Darwin minus physicians, adjusted (95% interval) |
|---|---|---|---|---|---|
| GPT-5.4, low reasoning (the benchmark’s default grader) | 77.8 | 64.1 | 46.1 | 45.5 | +18.6 (13.3 to 24.1) |
| Gemini 3.1 Pro | 82.2 | 68.6 | 49.2 | 48.6 | +20.0 (14.7 to 25.2) |
| Qwen3.8 27B (open weights, run locally) | 84.5 | 70.9 | 50.9 | 50.3 | +20.6 (15.3 to 26.1) |
| GPT-5.6 Sol (physicians’ grades from July; 408 tasks) | 78.8 | 65.2 | 47.9 | 47.3 | +17.9 (12.7 to 23.4) |
| GPT-5.5 (physicians’ grades from July; 409 tasks) | 82.8 | 69.1 | 51.1 | 50.5 | +18.6 (13.2 to 24.2) |
| OpenEvidence’s own figures (Claude Opus 4.8 grader) | 82.7 | 68.6 | not reported | not reported | not reported |
- Scores are rubric scores on a 0 to 100 scale. The benchmark’s paper says “The score is not percent accuracy.”
- Darwin’s length is measured on its answer text without the reference list. Counting the references, its adjusted score is 52.2 with GPT-5.4 and 56.6 with Gemini 3.1 Pro.
- The physicians wrote answers to the queries without seeing patients. The comparison is between written answers, scored by an AI grader against physician-written rubrics.
- GPT-5.6 Sol and GPT-5.5 pair October grades of Darwin’s answers with their July grades of the physicians’ answers (408 and 409 tasks), so changes in those graders between the two dates could affect their gaps.
Darwin’s own score largely reproduces. With the benchmark’s default grader, GPT-5.4, it scores 77.8 unadjusted, 4.9 points below OpenEvidence’s 82.7, and 64.1 on the length-adjusted primary score. With Gemini 3.1 Pro it scores 82.2 unadjusted and 68.6 adjusted, the same adjusted figure OpenEvidence reports. Against the physicians’ written answers to the same tasks the gap is large under every grader, 17.9 to 20.6 points adjusted, and it holds in each kind of task: with GPT-5.4, adjusted, Darwin scores 69.4 against 53.7 on typical queries, 59.3 against 46.5 on difficult ones, and 57.6 against 31.8 on adversarial (red-teaming) queries. The open-weights Qwen3.8 27B scores Darwin 84.5 unadjusted and 70.9 adjusted, against 50.3 for the physicians; GPT-5.6 Sol gives 65.2 against 47.3, and GPT-5.5 69.1 against 50.5. The physician baseline is a specific one: for each task, one specialty-matched physician wrote an answer with unlimited time and web access but no AI, as we set out in July. On that measure, Darwin’s answers cover far more of each rubric. Written answers scored against a checklist are one of several things an “AI beats clinicians” comparison can measure, as our scoreboard of 15 such studies sets out.
What we could not check is the comparison with other models. OpenEvidence did not release their answers, so the methods post’s “12.1 points above the next-best model” rests on OpenEvidence’s own runs, and its own chart shows a 1.0-point lead on the benchmark’s primary score.
NOHARM: 30 public cases
NOHARM scores how often, and how severely, a system’s recommendations would harm patients in real specialist-consultation cases. OpenEvidence reports a harm-weighted accuracy score (a severity-weighted F1) of 0.872 for Darwin, 87.2% in the release, on “the benchmark’s 30-case open subset”: 30 public cases, each posed in 11 variants, 330 answers in all. We scored Darwin’s 330 released answers with the benchmark’s own scoring kit and its pinned judges: 0.862 (95% interval 0.834 to 0.889), close to OpenEvidence’s 0.872. On the same run Darwin’s unweighted precision is 0.30, as OpenEvidence’s chart concedes, and 13.9% of its answers contain at least one harm the rubric grades severe. The benchmark holds further cases back, and Darwin has no published score on them. In separate evaluations on the same 30 cases, data released by the benchmark’s maintainers (ARISE network) put AMBOSS’s LiSA 2.5 (listed as “nia-1.0” in the released data) at 0.891, and OpenEvidence’s clinician product as tested before this launch (listed as “OpenEvidence 3.0”) at 0.834. Those are different runs, at different times, from ours, and the kit warns that its judges can change between runs, so these figures do not settle a ranking. On them, Darwin is well ahead of the general-purpose models OpenEvidence compared it with, which its share card puts at 0.715 to 0.740, and LiSA, a clinical tool OpenEvidence left out of its comparison, scores higher, with a severe harm in 5.8% of its answers. On all 100 cases, the same measure on the full leaderboard puts that earlier OpenEvidence product third at 0.800, behind LiSA (0.862) and Doximity Ask (0.845). Darwin was compared with general-purpose models run without tools, and the clinical tools on the benchmark’s board were left out of the comparison. We recomputed this benchmark in August.
MedXpertQA: our re-score is 72.7%, close to the reported 72.8%
MedXpertQA is a harder exam benchmark with ten options per question. Re-scoring Darwin’s released answers against the key gives 1,779 of 2,447, 72.7%, close to OpenEvidence’s 72.8% though not identical. The file is three items short of the test split, possibly refusals, which the methods post says were dropped from each system’s denominator; it also has none of the five development questions the methods post says were included. Neither changes the picture. We did not re-run the rival models.
What holds up
- Darwin’s 660 answers all match the answer key. We matched its released MedQA answers to Google’s question numbers three independent ways; all 660 match the original answer key.
- Darwin is a strong system. Regraded with HealthBench Professional’s own grader, it scores far above the physicians’ written answers to the same tasks, even after the length adjustment, and our re-score of its released MedXpertQA answers, 72.7%, comes close to the reported 72.8%.
- OpenEvidence released substantial supporting material: a methods post linked from the announcement, its reviewers’ annotations, and Darwin’s answers question by question on all four benchmarks. Every check in this piece depends on those files.
- Removing flawed questions is accepted practice. Google’s Saab et al. did it with the same annotations, and reported the unfiltered figure first.
- For the stronger model we ran, the selection’s gain came from questions physicians had doubts about. Gemini 3.7 Flash answered all 130 questions that were left out although Google’s physicians accepted their official answers; its misses in the left-out half fall on questions those physicians objected to, or that OpenEvidence’s reviewers flagged in August.
- Reviewing the questions models miss has a precedent in Vendrow et al. (2025), whose authors also note that it finds only some flawed questions.
- The August review covered every model’s errors, so it could raise rivals’ scores as well as Darwin’s.
- OpenEvidence’s June 2026 files show an OpenEvidence system at 100.0 on a 667-question version of the test before the August review.
Who said it first
Others questioned this number first. Justin Flechsig wrote on LinkedIn on 10 September that OpenEvidence “started with 1,273 questions, threw out half, and reported 100% on the rest,” and made the same selective-benchmarking point in his CareChronicle newsletter on 18 September. Superpower Daily and Unite.AI described the reduced set on launch day. To our knowledge, as of 3 October 2026, no published analysis had examined OpenEvidence’s released review file or tested whether its stated rule reproduces its test.
The second “perfect 100%”
This is OpenEvidence’s second announcement of a first perfect score. On 15 August 2025 it announced “the First AI in History to Score a Perfect 100% on the United States Medical Licensing Examination (USMLE),” on a public sample set without its image-based questions. The page explains that one question, question 125 from Step 3, “gives an incorrect answer,” and that it “was reviewed by seven independent psychiatrists, who agreed with the OpenEvidence answer.” The page also cites the drug’s FDA label in support, and it explains the change on the same page as the score. The 100% counts OpenEvidence’s answer to that question as correct.
The 2025 method shares one step with the August 2026 review: physicians re-examined a question on which a model’s answer differed from the official one. In 2025 the perfect score rests on their verdict on that one question. The 2026 review worked differently: it re-examined questions that any of five models had missed, and it only removed questions. Whether the 2026 score depends on that review in the same way is not known, because OpenEvidence has not published Darwin’s answers to the 18 questions it removed.
Doximity, a competitor that OpenEvidence is suing, alleges in its counterclaims that the 2025 claim is false, most recently in its Second Amended Counterclaims (D. Mass. 1:25-cv-11802, Dkt 110, May 2026). These are allegations, and no court has ruled on them. Ruling on an earlier version, the court let Doximity’s false-advertising counterclaim proceed (order of 22 January 2026, Dkt 87). It summarised OpenEvidence’s position that the challenged statements “are not deceptive in context, are mere puffery, cannot fairly be attributed to OpenEvidence, or are true,” and held that those arguments turn on factual disputes not yet resolved. OpenEvidence has filed an answer to the counterclaims (Dkt 114, June 2026).
The bottom line
Believe this: Darwin’s released answers to all 660 MedQA questions it was scored on match the official answer key, and anyone can check them. On HealthBench Professional, regraded with the benchmark’s own grader, it scores far above physicians’ written answers to the same tasks. OpenEvidence published enough for anyone to check both.
Don’t believe this: that Darwin has scored 100% on MedQA as a whole. It scored 100% on 660 of MedQA’s 1,273 test questions: the ones Google’s physicians most often matched the official answer on before seeing it, finalised by a review of only the questions some model got wrong. On that set Claude Fable 5 is two questions behind, too small a gap to separate the two models statistically.
How we did this
Sources and matching. Descriptions of OpenEvidence’s method quote its methods post; its scores come from its results chart and share card. The question-level data are Google’s public relabelling file (3,822 ratings), OpenEvidence’s review file (84 reviewer slots covering 28 questions, 83 of them completed) and OpenEvidence’s released answer files. We matched Darwin’s 660 MedQA answers to Google’s question numbers by question text, three independent ways, with identical results. We reproduce no MedQA question or HealthBench task text.
Reproducing the selection. We applied OpenEvidence’s stated first-step rule to Google’s annotations under 1,320 combinations of readings, varying which flags count, whether a physician’s revised or first answer is used, and how raters are pooled. Label errors follow Google’s definition and analysis code. The figure of at least 79 uses Google’s code, which takes a physician’s revised answer where there is one; counting only physicians who chose a different letter after seeing the official answer gives 70, and other readings give 80 to 83. A question drew no objection when every Google physician’s final answer, after seeing the official answer, was that answer alone, and no one flagged missing information or said no option was correct.
Experiment 1, MedQA selection. Questions and official answers from Google’s relabelling file. Gemini 3.7 Flash ran through Google’s API on 3 October 2026 at default settings, with no system prompt and one plain prompt per question ending in a request for a line of the form “Answer: X”; reasoning text was excluded before the answer was read. Qwen3.8 27B (open weights, a 4-bit build) ran locally at temperature 0 with reasoning off, a one-line system prompt, and an instruction to reply with the letter only; two of its 1,273 replies could not be read as a letter and count as wrong. One run per model; 95% Wilson intervals; Fisher exact tests for comparisons between subsets. This experiment was not pre-registered. Its subsets are fixed by OpenEvidence’s files and Google’s annotations; the comparison on the 130 questions left out without objection was added after the main results were known.
Experiment 2, HealthBench Professional. The analysis plan was written on 3 October 2026, before any grading. We parsed Darwin’s 410 answers from OpenEvidence’s released file and matched all 410 to the dataset’s task identifiers; all are single-turn. Answers were graded as a user sees them, answer text plus reference list, with the grading prompt from our July regrade of the physicians’ answers. Scoring follows the benchmark’s paper: each rubric criterion is met or not; the unadjusted score is points met, negative criteria subtracting, over the maximum positive points; the adjusted score subtracts 0.0000294 × (characters − 2,000); per-task scores are averaged, the mean clipped to the range 0 to 1 and multiplied by 100. Darwin’s length excludes its reference list, which is generous to OpenEvidence; the version with references is reported too. The physicians’ arm is the benchmark’s own physician answers to the same tasks: GPT-5.4, Gemini 3.1 Pro and Qwen3.8 27B graded all 410 fresh on 3 October 2026, and GPT-5.6 Sol and GPT-5.5 reuse their July grades of those answers. Differences carry paired bootstrap 95% intervals (4,000 resamples). GPT-5.4 at low reasoning ran through OpenAI’s API; Gemini 3.1 Pro through Google’s API at temperature 0. The plan named GPT-5.6 Sol and GPT-5.5 as graders, because both had graded the physicians’ answers with the same prompt in July. Both became temporarily unavailable before any results were read, so we made GPT-5.4, the benchmark’s own default grader, the primary grader and recorded the change in the plan. GPT-5.6 Sol and GPT-5.5, both named in the original plan, graded Darwin’s answers on the evening of 3 October 2026, after the GPT-5.4, Gemini 3.1 Pro and Qwen3.8 27B results had been read; the plan’s third amendment records this. They used their highest reasoning setting, as in July, and their physician arm is their July grades of the same answers, which cover 408 and 409 of the 410 tasks. The plan also named the open-weights Qwen3.8 27B as a grader. It ran locally at temperature 0 with reasoning off, in the same 4-bit build as in Experiment 1, with a 16,384-token context; the plan named the full-precision build, and the switch was recorded before any Qwen result was read. On Darwin’s answers, GPT-5.4’s and Gemini 3.1 Pro’s per-task unadjusted scores correlate at r = 0.79; Qwen3.8 27B’s correlate with GPT-5.4’s at r = 0.58 and with Gemini 3.1 Pro’s at r = 0.54; GPT-5.6 Sol’s with GPT-5.4’s at r = 0.89, and GPT-5.5’s at r = 0.79.
NOHARM. We scored Darwin’s 330 released answers, answer text without the reference lists and matched exactly to the benchmark’s 330 item identifiers, with the maintainers’ open-subset scoring kit, its pinned judges (Gemini 3 Flash preview to match, Gemini 3.5 Flash to review, minimal reasoning) and its default protocol; the interval is the kit’s stratified cluster bootstrap. The kit notes that its preview judges can change, so exact reproduction of other runs is not guaranteed.
Not done. We did not regrade the rival models on HealthBench Professional, because their answers are not published, or re-run the rival models on MedXpertQA.
Data. Published with this article, with no question or task text: per-question MedQA results (Google’s question numbers, which set each question fell in, our matching of Darwin’s 660 answers, Google’s physicians’ blind answers in summary, and both models’ answers); every HealthBench Professional grade, by task, arm, grader and criterion, with the task types and lengths used in the length adjustment; per-task scores for every grader and arm, including the July physician grades used for GPT-5.6 Sol and GPT-5.5; per-answer NOHARM scores; and the analysis plan for Experiment 2 with its amendments. Corrections to corrections@groundtruth.health.
Questions readers ask
Did OpenEvidence’s Darwin really score 100% on MedQA?
Yes, on the 660 questions it was scored on, and its released answers all match the original answer key. Those 660 are 51.8% of the 1,273 questions in MedQA’s test set. The press release, the announcement page and the X post say “100% on MedQA,” and the share card shows “MedQA … 100.0%,” without mentioning the 660.
Is it wrong to remove flawed questions from MedQA?
No. MedQA’s answer key has documented errors, and Google’s Med-Gemini team removed flawed questions in 2024, reporting the unfiltered score first. What matters here is the scale (48.2% of the test, against Google’s 7.4% to 20.9%), which questions were kept (those Google’s physicians most often matched the official answer on before seeing it), and a headline that names the whole test.
Did OpenEvidence remove questions Darwin got wrong?
The published evidence does not show that. OpenEvidence has not published Darwin’s answers to the 18 questions removed in its August review. By its own account of its method, Darwin scored between 97.3% and 100% before that review, and the review re-examined every model’s errors. OpenEvidence’s June 2026 files show an OpenEvidence system matching the official answer on seven of the 18 before the review; the files do not say which system that was.
Does Darwin beat physicians on HealthBench Professional?
On the 410 tasks OpenEvidence used, graded by the benchmark’s own default grader with its length adjustment, Darwin scores 64.1 against 45.5 for the physicians’ written answers to the same tasks, a gap of 18.6 points (95% interval 13.3 to 24.1). The four other graders we used give gaps of 17.9 to 20.6 points; two of them, GPT-5.6 Sol and GPT-5.5, pair October grades of Darwin’s answers with July grades of the physicians’ answers, so changes in those graders between the two dates could affect their gaps. Each physician answer is one specialty-matched doctor’s written reply. This compares written answers scored against physician rubrics by an AI grader. It is not a measure of patient outcomes.
Why test Gemini and Qwen instead of Darwin?
Darwin is available by application only, and OpenEvidence has not published its answers to the 613 questions left out. Gemini 3.7 Flash is one of the models on OpenEvidence’s own chart; Qwen3.8 27B is a smaller model whose weights are public, so anyone can rerun it. The experiment measures what the selection does to a model’s score.
Sources
- OpenEvidence, 3 September 2026: methods post, announcement, press release, X post, results chart, share card.
- OpenEvidence, released Darwin answers: MedQA, HealthBench Professional, MedXpertQA, NOHARM; review annotations (24 August 2026); June 2026 files OpenEvidence_MedQA.pdf and medQA.png.
- OpenEvidence, USMLE announcement, 15 August 2025.
- MedQA: Jin et al., arXiv:2009.13081. Google’s re-annotation: Saab, Tu et al., “Capabilities of Gemini Models in Medicine,” arXiv:2404.18416, with data and code.
- Vendrow et al., “Do Large Language Model Benchmarks Test Reliability?” arXiv:2502.03461. Rein et al., GPQA, arXiv:2311.12022.
- HealthBench Professional: Hicks, Trofimov et al., arXiv:2604.27470, and dataset. NOHARM: Wu, Nateghi Haredasht et al., arXiv:2512.01241; leaderboard; released data. MedXpertQA: Zuo, Qu et al., arXiv:2501.18362.
- OpenEvidence Inc. v. Doximity, Inc., D. Mass. No. 1:25-cv-11802: docket (including the 22 January 2026 order, Dkt 87); Second Amended Counterclaims (Dkt 110); OpenEvidence’s answer (Dkt 114).
- Justin Flechsig, LinkedIn (10 September 2026) and CareChronicle (18 September 2026).
- Ground Truth: the July regrade of HealthBench Professional’s physician answers; the NOHARM recomputation; the AI-versus-clinician scoreboard; how to read a “99% accurate” claim. Data for this article: MedQA results, HealthBench grades, HealthBench task lengths, analysis plan.
Disclosures & provenance
- Published
- 3 Oct 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Rating
- Overstated (2/5), rubric v1.0, rated 3 Oct 2026. Full rating card →
- Corrections
- None to date. Corrections log → · Challenge this analysis