On 21 June 2026, News-Medical.net ran the headline “Google’s AMIE beats doctors on key simulated disease-management tasks.” Three days earlier Digit.in had gone with “AI is beating human doctors in two key areas, and it’s a good thing,” and reported that specialists rated AMIE’s treatment recommendations as sufficiently precise in 96% of cases at the first visit, against 62% for the physicians. Both stories rest on a real paper: “Towards conversational artificial intelligence for disease management,” by Valentin Liévin, Anil Palepu and colleagues at Google, published in Nature on 17 June 2026, open access. The 96% and the 62% are in it. So is the sentence that travelled furthest, from the paper’s own Discussion: “There were no domains in which PCPs outperformed AMIE.” Nobody invented anything.

Two months later, on 11 August 2026, Google Research went further in its own voice: “we present the first demonstration of an AI system exhibiting expert-level performance in real-time clinical video consultations.” That sentence is the claim we rate, and we rate it Overstated (2/5). The preprint behind it never defines “expert-level.” The experts it was measured against were ten US primary care physicians averaging four and a half years since residency, asked to keep their cameras off throughout, scored by an unblinded panel averaging nearly fifteen. The results underneath are real, and in places strong. The standard is thinner than the word implies.

The June story is the same story one rung earlier. The paper’s own word for its main result is “non-inferior,” and no non-inferiority margin is defined anywhere in it or its supplement. The 96% is a rating proportion: the share of scripted scenarios in which three specialists gave a favourable median rating on one of fifteen axes, the one scoring how precisely a plan names the drug, dose, duration and route. The physicians were 21 board-certified primary care physicians in Canada and India, typing to actors over text chat, judged on scenarios written to UK guidance they had not necessarily practised under, at intervals of about two days where the scenarios describe weeks. The paper says all of this, and says the system is not ready for clinical care. The coverage carried the verb and the number.

In fairness to News-Medical, whose headline we quote above: its body carries “non-inferior,” the patient actors, the text chat and the UK-guideline design. The escalation there is confined to the headline. That is why the rated claim is Google’s own August sentence rather than a press headline — the August claim is first-party, and nothing downstream of it walks it back.

Nature published something else with the paper: the peer-review file. Four rounds of referee reports and author responses, 91 pages, open access, one click from the article. A physician referee wrote that the physician baseline “seems highly questionable.” A second asked whether AMIE was “already supplanted simply by newer versions of the base LLM.” The authors answered that if the study “were repeated today we might choose another approach.” We ran that referee’s question as an experiment, and the answer is in the middle of this piece.

The steel-man first: this is the most rigorous study of its kind

If a reader finishes this piece thinking Ground Truth called it a bad study, the piece has failed. The June paper is, on the record, the most careful longitudinal comparison of an AI system against clinicians yet published. It is randomised and blinded, running 100 scenarios across five specialties, each over three visits, each visit run twice: once with AMIE, once with a physician. The 21 physicians were board-certified, median nine years post-residency. Thirty specialist raters, matched by specialty and jurisdiction, rated every consultation in triplicate, and the 45 statistical comparisons were corrected for multiple testing. Both arms could draw on the same 627 guideline documents, 10.5 million tokens of NICE and BMJ Best Practice material. The authors chose two conditions they argue favour the humans: visit intervals of a day or two where the scenarios describe weeks, which they say “probably increased human performance,” and unlimited guideline access for the physicians after each consultation.

AMIE’s plans scored at least as well as the physicians’ on all fifteen axes and better on the specificity and guideline-grounding axes. That is a real result. The paper states plainly that its findings “do not suggest that AMIE is ready for clinical care.” Everything a reader needed in order not to be misled was already on the page. What follows is about what happened to it.

What “96% versus 62%” counts

The metric behind every headline number in this paper is a rating proportion: the share of scenarios in which the median of three specialist ratings was favourable. It measures neither diagnostic accuracy nor any patient outcome. Nine of the fifteen axes are yes/no questions; the rest are Likert scales binarised at their top two options. The 96% is the treatment-preciseness axis, which the paper defines as naming “the precise antibiotic, its dose, the duration of treatment and administration route.” It is a specificity-of-instructions measure, and a language model that emits a structured plan is favoured on it close to by construction. It is also visit 1 only. At visits 2 and 3 the same axis reads 95% against 65% and 95% against 67%.

The number that travelled: treatment preciseness, favourable specialist ratings, by visit

020406080100AMIE, visit 1AMIE, visit 1: 96 · the headline 96%96Physicians, visit 1Physicians, visit 1: 62 · the headline 62%62AMIE, visit 2AMIE, visit 2: 9595Physicians, visit 2Physicians, visit 2: 6565AMIE, visit 3AMIE, visit 3: 9595Physicians, visit 3Physicians, visit 3: 6767Scenarios rated favourable by specialists (%)
Source: Liévin, Palepu et al., Nature (17 June 2026), Results, treatment preciseness axis; all three visits P < 0.001.
View chart data
The number that travelled: treatment preciseness, favourable specialist ratings, by visit data table
ItemScenarios rated favourable by specialists (%)Note
AMIE, visit 196the headline 96%
Physicians, visit 162the headline 62%
AMIE, visit 295
Physicians, visit 265
AMIE, visit 395
Physicians, visit 367
Read before citing
  • Preciseness means naming the precise antibiotic, dose, duration and route: a specificity-of-instructions axis on which a system that emits a structured plan is favoured close to by construction. It is not diagnostic accuracy and not a patient outcome.
  • The metric on every axis is the proportion of scenarios whose median rating from three specialists was favourable, with 9 of 15 axes yes/no and the rest binarized at the top two options of their scale.
  • The physicians were 21 board-certified primary care physicians in Canada and India, median nine years post-residency, judged on scenarios built to UK guidance; the paper concedes their familiarity with those guidelines may have been limited.

The second omission sits in the same sentence. “Non-inferior” appears five times in the paper; a non-inferiority margin appears nowhere. Across the paper and its supplement the nouns “non-inferiority” and “noninferiority” occur zero times, and “margin” survives only as “marginally,” describing small differences, and in the title of a cited statistics paper. The word carries a regulatory meaning this design never claimed to earn: the paper uses it descriptively, to say that AMIE’s plans scored at least as well as the physicians’ across the fifteen axes. The press then upgraded a descriptive “non-inferior” to “beats.”

The file nobody opened

The paper was received on 17 March 2025 and accepted on 4 June 2026: 444 days, four rounds. Nature published the referee reports and author responses alongside it, open access. It is the most informative document in this story, and none of the coverage we coded cites it.

Referee 1, a physician, opened on the human baseline. Reading the example consultations, they found the physician in the first scenario “highly questionable in how they are acting,” and reported showing it to other physicians “who agree that it’s concerning what the ‘baseline’ for physicians were” (peer-review file, p.4). Having used the system, the same referee found it “reminding me of a medical student without diagnostic abilities” (p.5). In the final round: “I would not use or suggest implementation of such a tool without more transparency” (p.10).

That is the question this publication asks of every AI-versus-clinician study — was the human baseline matched? — asked by a referee, on the record, in a file attached to the paper. Referee 3 asked the other half of it: the guidelines “appears to be entirely from the UK but the 21 PCPs who participated in the study are based in India and Canada — is this a fair comparison” (p.4). The authors conceded the point and wrote it into the published paper, which now states that the physicians’ familiarity with those guidelines “may have been limited.”

One side of the comparison is public. The other is not.

AMIE’s consultations for all 120 scenarios, the 100 in the main study and 20 more, were released: a 2,483-page PDF and a CSV. The physicians’ consultations were not. The authors explained why, in the file (p.87): releasing the physician outputs “is infeasible in our particular setting as it would require additional contracting with each of the 20+ clinicians involved, to clear permissions for data release.” The reason given is a permissions problem, and the same paragraph commits to releasing everything on AMIE’s side. Still, the asymmetry deserves stating plainly: in a study whose entire headline is a comparison between two arms, the AI’s transcripts are downloadable and the doctors’ are not. Output for two sample scenarios was shared. The referee who called the baseline questionable saw two examples, and a reader today can see the same two.

What the authors actually claimed, in their own words

Two passages from the authors’ responses settle what the paper claims and does not. On the architecture (p.76): the ablation results support the benefit of their design on the base model available at the time, “rather than superiority of the specific agentic architecture compared to other models, which is not an explicit claim we aim to make.” On readiness (p.85): “the system investigated in this research is not ready for real-world translation… This research is an art-of-the-possible demonstration.” The authors told their referees they were not claiming superiority. The press reported superiority.

Peer review worked here. That is the point, not a gotcha.

Peer review demonstrably improved this paper. The original submission used a single rater per case; triplicate specialist rating, the basis of every published number, was added only after a referee asked about inter-rater agreement (p.24). A referee spot-checked an implausible p-value; the authors investigated, found a defect in a pandas.crosstab call that affected 5 of the 45 McNemar tests where a contingency-table row was all zeros, and tabulated every affected result (pp.37–38). Every headline number now in circulation exists in its published form because reviewers pushed. That process made the published numbers more trustworthy.

The fix reached the paper’s figure but not all of its prose. In Figure 4, the visit-2 “follow-up appropriate” cell carries no significance marker. The Results text still reports that same comparison — “providing appropriate follow-up recommendations (100% versus 97%, P < 0.001, for visit 2)” — at a value consistent with the pre-correction calculation. The arithmetic is checkable from the published paper alone. Across 100 paired scenarios, 100% versus 97% means three discordant pairs, and an exact McNemar test with three discordant pairs cannot return anything below 0.25. That is the corrected raw value in the authors’ own table, where the cell is marked not significant. A second slip sits in the same paragraph, pointing the other way: treatment alignment with guidelines at visit 1 is printed as 91% versus 87%, where the table gives 91% versus 78% at the same p-value, and visits 2 and 3 match exactly. On 1 September 2026 we checked the live version of record, dated 22 July. Both sentences stand, with no correction notice attached.

AMIE still scored higher on the follow-up measure. What the correction removes is the claim that a three-point gap is statistically distinguishable; the gap itself stands. The two errors run in opposite directions — one flatters the system, the other understates it — which is the signature of a proofing failure. This is one of 45 tests, on an axis where both arms score at or above 97%, and it changes none of the paper’s conclusions. The authors found the defect themselves, disclosed all five affected results, and correctly dropped the other newly non-significant result from the text. It is visible at all only because they documented their own correction.

The question the referees left open

Referee 3 asked it directly (p.10): does the ablation not “call into question several central claims,” and “Is AMIE already supplanted simply by newer versions of the base LLM?” Referee 2, who signed the review, observed that later base models without AMIE often outperform earlier models with it, which “may somewhat diminish the perceived contribution of AMIE itself.” The authors’ answer is the hinge of this piece (p.57): their agent design “was just one configuration; if this study were repeated today we might choose another approach based on performance of the stronger base model.” The rest of that passage matters too. The agentic approach, they say, “showed meaningful gains over the corresponding Gemini 1.5 base model which was available at the time of the study.” Their point is that its value was indexed to a base model that has since moved.

The published paper says the same in its own results. AMIE’s backbone was Gemini 1.5 Flash. The paper reports that contemporary models including Gemini 2.5 Pro, o3 and GPT-5 achieved performance “largely comparable” with AMIE on its drug-knowledge benchmark, that Gemini 2.5 Flash alone “substantially improv[ed] conversational performance” over it, and that the agent scaffolding gave “little to no improvement” on the newer base model. Three of the 17 articles we coded carried that finding.

So we repeated it today

Read this section as a benchmark floor. It does not retest the simulated disease-management result, and nothing here bears on 96% versus 62%. It tests the separate, public benchmark the same authors built and released with the paper, so the referee’s question gets an empirical answer.

RxQA is 300 multiple-choice questions on drug knowledge, generated from OpenFDA labels, revised by pharmacists, and published by Google under an Apache 2.0 licence. A further 300 questions built from the British National Formulary are withheld for licensing. We ran the 300 public questions with the prompt from the repository’s own notebook, unmodified, at temperature 0, closed-book, matching the closed-book column of the paper’s Supplementary Table SI.34. The paper never states its own inference settings, so we ran each model twice, with default thinking and with thinking disabled and verified to produce zero thought tokens.

The harness validates on the one model both we and the authors measured. On Gemini 2.5 Flash our scores are 68.7% with thinking and 60.0% without, against the paper’s published 65.5; neither difference is significant (p = 0.41 and p = 0.16). Their number sits inside our bracket. That is what licenses the comparison that follows.

RxQA, 300 public OpenFDA questions, closed-book: the published AMIE score against a cheap general model run one year later

4050597080Six physicians (paper)Six physicians (paper): 48.3 (no interval published)48.3AMIE on 1.5 Flash (paper)AMIE on 1.5 Flash (paper): 59.0 (no interval published)59.02.5 Flash, bare (paper)2.5 Flash, bare (paper): 65.5 (no interval published)65.5AMIE on 2.5 Pro (paper)AMIE on 2.5 Pro (paper): 74.0 (no interval published)74.02.5 Flash, ours, no thinking2.5 Flash, ours, no thinking: 60.0 (95% CI 54.4 to 65.4; vs 65.5: p=0.16)2.5 Flash, ours, no thinking: 60.0 (95% CI 54.4 to 65.4; vs 65.5: p=0.16)60.0 [54.4, 65.4]2.5 Flash, ours, thinking2.5 Flash, ours, thinking: 68.7 (95% CI 63.2 to 73.7; vs 65.5: p=0.41)2.5 Flash, ours, thinking: 68.7 (95% CI 63.2 to 73.7; vs 65.5: p=0.41)68.7 [63.2, 73.7]3.6 Flash, ours, no thinking3.6 Flash, ours, no thinking: 71.7 (95% CI 66.3 to 76.5; vs 59.0: p=0.001)3.6 Flash, ours, no thinking: 71.7 (95% CI 66.3 to 76.5; vs 59.0: p=0.001)71.7 [66.3, 76.5]3.6 Flash, ours, thinking3.6 Flash, ours, thinking: 74.0 (95% CI 68.8 to 78.6; vs 59.0: p<0.001)3.6 Flash, ours, thinking: 74.0 (95% CI 68.8 to 78.6; vs 59.0: p<0.001)74.0 [68.8, 78.6]3.6 Flash, options shuffled3.6 Flash, options shuffled: 73.0 (95% CI 67.7 to 77.7; vs 74.0: p=0.74)3.6 Flash, options shuffled: 73.0 (95% CI 67.7 to 77.7; vs 74.0: p=0.74)73.0 [67.7, 77.7]Accuracy (%) on 300 questions; 95% intervals for our runs
Source: Published anchors: Liévin, Palepu et al., “Towards conversational artificial intelligence for disease management,” Nature (17 June 2026), Supplementary Table SI.34, OpenFDA subset, closed-book column. Our runs: Ground Truth rerun of github.com/Google-Health/rxqa (Apache 2.0), repo prompt unmodified, temperature 0, closed-book, 23 July 2026; raw responses and scorer published with this article.
View chart data
RxQA, 300 public OpenFDA questions, closed-book: the published AMIE score against a cheap general model run one year later data table
Run (full description in the Note column)Accuracy (%)95% intervalInterval typepSig.Setting
Six physicians (paper)48.3not reportedno interval publishedyesSix primary-care physicians (paper, SI.34) — published mean, no interval reported
AMIE on 1.5 Flash (paper)59.0not reportedno interval publishedyesAMIE on Gemini 1.5 Flash (paper, SI.34) — the specialized system as published
2.5 Flash, bare (paper)65.5not reportedno interval publishedyesGemini 2.5 Flash, bare model (paper, SI.34) — the paper's own ablation
AMIE on 2.5 Pro (paper)74.0not reportedno interval publishedyesAMIE on Gemini 2.5 Pro (paper, SI.34) — the paper's best configuration
2.5 Flash, ours, no thinking60.054.4 to 65.495% CIvs 65.5: p=0.16yesGemini 2.5 Flash, our rerun, thinking off — vs paper's 65.5: p = 0.16
2.5 Flash, ours, thinking68.763.2 to 73.795% CIvs 65.5: p=0.41yesGemini 2.5 Flash, our rerun, default thinking — vs paper's 65.5: p = 0.41
3.6 Flash, ours, no thinking71.766.3 to 76.595% CIvs 59.0: p=0.001yesGemini 3.6 Flash, our rerun, thinking off — vs AMIE 59.0: +12.7, p = 0.001
3.6 Flash, ours, thinking74.068.8 to 78.695% CIvs 59.0: p<0.001yesGemini 3.6 Flash, our rerun, default thinking — vs AMIE 59.0: +15.0, p < 0.001
3.6 Flash, options shuffled73.067.7 to 77.795% CIvs 74.0: p=0.74yesGemini 3.6 Flash, answer options shuffled — contamination probe; paired exact McNemar p = 0.74 vs unshuffled
Read before citing
  • This is a benchmark-floor sidebar, not a retest of the headline. The “96% versus 62%” result comes from the simulated multi-visit consultation study, not from RxQA. Nothing in this chart bears on that number.
  • No direct head-to-head with AMIE is possible: AMIE was never released, and Gemini 1.5 Flash, its backbone, has been retired from Google's API. The diamonds are the paper's published figures; the circles are our runs. The paper does not state its inference settings, so we report both settings.
  • Our harness reproduces the one model both sides measured: Gemini 2.5 Flash scores 68.7 (thinking on) and 60.0 (thinking off) against the paper's 65.5, neither difference significant.
  • Closed-book, OpenFDA half only. The 300 BNF questions are withheld for licensing, and the per-question medication labels needed for the open-book condition were not released.
  • The published answer key is skewed: option C is correct for 113 of 300 questions. Shuffling the options removes any position advantage; the score held (73.0 vs 74.0, paired exact McNemar p = 0.74).
  • Published anchors are printed without intervals because the paper reports none; the six-physician baseline is a mean across physicians, and the paper itself says it indicates no measure of real-world competence.

Gemini 3.6 Flash, a cheap general-purpose model with no medical scaffolding, no retrieval and no agent wrapper, scores 74.0% (95% interval 68.8–78.6). The full AMIE system, as published, scored 59.0. The six-physician baseline the paper used for this benchmark scored 48.3. Against the published AMIE figure that is 15.0 points, p < 0.001; on the conservative thinking-off setting, 71.7%, still 12.7 points, p = 0.001. It also matches the paper’s own best configuration, AMIE on Gemini 2.5 Pro at 74.0, without the scaffold.

Two controls, raised before anyone else raises them. First, contamination: RxQA’s answers have been on GitHub since March 2025, so a model might have memorised the letters. We permuted the options deterministically and recomputed the key: 74.0% became 73.0%, paired exact McNemar p = 0.74. The published key is also skewed, with option C correct for 113 of 300 questions, so shuffling strips any guess-C advantage; the score held. Second, inference settings. Thinking moves Gemini 2.5 Flash by 8.7 points (p = 0.0025) and Gemini 3.6 Flash by 2.3 (p = 0.32). On the older model, a setting the paper never reports moves the score by more than its entire published gain of AMIE over its base model.

The setting the paper never states moves the older model by nine points

020406080100Gemini 2.5 Flash, thinking onGemini 2.5 Flash, thinking on: 68.7 · paired vs thinking off: +8.7 points, p = 0.002568.7Gemini 2.5 Flash, thinking offGemini 2.5 Flash, thinking off: 60 · thinkingBudget = 0, verified 0 thought tokens60Gemini 3.6 Flash, thinking onGemini 3.6 Flash, thinking on: 74 · paired vs thinking off: +2.3 points, p = 0.32, not significant74Gemini 3.6 Flash, thinking offGemini 3.6 Flash, thinking off: 71.7 · thinkingLevel = minimal, verified 0 thought tokens71.7RxQA closed-book accuracy (%), n = 300
Source: Ground Truth rerun of RxQA (github.com/Google-Health/rxqa), 300 OpenFDA questions, closed-book, temperature 0, 23 July 2026. Paired comparisons by exact McNemar test on the same 300 questions.
View chart data
The setting the paper never states moves the older model by nine points data table
ItemRxQA closed-book accuracy (%), n = 300Note
Gemini 2.5 Flash, thinking on68.7paired vs thinking off: +8.7 points, p = 0.0025
Gemini 2.5 Flash, thinking off60thinkingBudget = 0, verified 0 thought tokens
Gemini 3.6 Flash, thinking on74paired vs thinking off: +2.3 points, p = 0.32, not significant
Gemini 3.6 Flash, thinking off71.7thinkingLevel = minimal, verified 0 thought tokens
Read before citing
  • The Nature paper does not report whether its own Gemini runs used thinking. On the older model, that one unstated setting moves the score by more than the paper's entire published gain of AMIE over its base model (59.0 vs 48.1, 10.9 points).
  • Same shape as the paper's own ablation, one generation later: the scaffolding mattered for the older base model and stopped mattering for the newer one.
  • Our numbers are not directly comparable to the paper's absolute scores (different harness, unknown grading details on their side); the within-chart comparisons are paired on identical questions.

What our own test cannot show

The study that did not travel

The same lab ran AMIE on real patients. The feasibility study, arXiv 2603.08448, was prospective, pre-registered and single-arm, at Beth Israel Deaconess Medical Center: 100 patients completed an AMIE interaction, 98 completed the physician visit, and the comparisons are on those 98. There, the physicians outperformed AMIE on the practicality of management plans (p = 0.003) and on cost-effectiveness (p = 0.004), with no significant difference on the differential diagnosis.

Three cautions against conflating the two. That study used a different AMIE, built on Gemini 2.5 Pro and switched mid-study to 2.5 Flash. Its comparator pool was mostly trainees: 61 residents, 11 attendings and five nurse practitioners. And its “90%” differential-diagnosis figure means the diagnosis appeared within the first seven ranked candidates in 88 of 98 cases; top-1 was 56%. The result that physicians beat the system on practicality and cost appears in none of the articles we coded.

The third rung: August, on video

On 11 August 2026, Google Research announced AMIE (Video), a real-time video-consultation configuration built on Gemini 3 Flash and Gemini 3.1 Pro, with a preprint, arXiv 2608.09861, that has not been peer-reviewed. The blog’s claim: “we present the first demonstration of an AI system exhibiting expert-level performance in real-time clinical video consultations.” The design is again a randomised simulated consultation study: 100 scenarios, 15 professional patient actors, 300 consultations across three arms, an independent panel of 20 physicians as raters. The headline result is a case-specific rubric score of 83% for AMIE (Video) against 68% for the physicians (p = 1.32 × 10−9, Table A.7). Top-1 diagnostic accuracy was 0.91 against 0.77 (p = 0.039); top-3 was 0.98 against 0.90, not significant (p = 0.179).

Three things about the comparator. The physician arm was ten US board-certified primary care physicians averaging 4.5 years post-residency, range one to nine, against evaluators averaging 14.95 years since residency. The physicians “were asked to turn their cameras off throughout the consultations,” to match an AI system that has no face; the preprint notes this “may have impacted their ability to establish rapport and empathy.” And the raters were not blind: full blinding “was not possible,” the preprint says, because AMIE’s synthetic voice “remains audibly distinct from human physician speech.”

One sentence in the blog compresses the patient-actor results past what the table supports. The blog says actors “rated AMIE (Video) favorably on empathy, rapport, and confidence in care compared to both PCPs and AMIE (Text).” Against AMIE’s own text version the direction holds. Against the physicians none of the three named dimensions reaches significance: Table A.9 gives building rapport 0.81 against 0.82 (adjusted p = 0.830) and partnership building 0.66 against 0.74 (p = 0.750), both leaning the physicians’ way, with confidence in care and empathy leaning AMIE’s but not significantly. Where the actors did significantly prefer AMIE was elsewhere, on assessing and explaining their condition (both p = 0.045). That is what the preprint’s abstract reports — actors “preferred AMIE’s approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building” — and it is in the blog’s own results figure. It is the blog’s prose that reaches instead for empathy, rapport and confidence.

Put the three AMIE studies side by side, as we have done across fifteen others, and the human comparator changes materially each time. The studies differ in modality, jurisdiction and rater design too, so this is not one benchmark being progressively weakened. It is a reason to read the comparator before the headline, every time.

The physician comparator in each AMIE study

The physician comparator in each AMIE study
StudyPhysicians in the comparison armSettingWho ratedHeadline
Diagnostic dialogue, Nature, April 202520 board-certified PCPs, 3–25 years post-residency, median 7; ten each from India and CanadaText chat, 159 scenarios, 20 actorsSpecialist physicians and actors“Superior performance on 30 out of 32 axes”
Disease management, Nature, June 202621 board-certified PCPs in Canada and India, median 9 years post-residencyText chat, 100 scenarios × 3 visits, actors, UK-guideline scenarios30 specialists, triplicate“Non-inferior”; 96% vs 62% preciseness
Video consultations, arXiv preprint, August 202610 US board-certified PCPs, mean 4.5 years post-residency, range 1–9, cameras offVideo, 100 scenarios, 15 actors20 PCPs, mean 14.95 years since residency, not blind to arm“Expert-level”; 83% vs 68% rubric score
Source: Tu, Schaekermann, Palepu et al., Nature (9 April 2025), Methods; Liévin, Palepu et al., Nature (17 June 2026), Methods; Nagda, Lee et al., arXiv:2608.09861 (11 August 2026), Sections A.6.5 and A.6.7.

The blog describes the study as involving “a group of 30 board-certified primary care physicians.” That is the total: ten in the consulting arm and twenty on the evaluation panel, which the preprint says was “entirely separate” from the consultation cohort. The number of physicians AMIE was compared against is ten.

None of this makes the August result wrong. AMIE (Video) was rated ahead of the physicians on most general criteria and far ahead on physical observation and guided examination, where the video modality is the point. But “expert-level” is a claim about a standard, and the standard on offer is the arm described above, scored by raters who could hear which arm they were listening to. The June paper’s referees had asked whether the baseline was matched. The August study did not answer that question; it changed the comparator.

What survived the trip

We collected the coverage of the June paper, coded each article for whether it carried each of seven conditions the result depends on, and on 1 September 2026 re-fetched every article and checked each quoted passage at its origin. Seventeen articles published between 17 and 30 June 2026 were reachable and re-verified; three that were reachable in July now return access errors and are excluded, along with two whose addresses we could not recover. The counts below are for those 17. The full coverage ledger lists all 39 items we examined, including every exclusion and its reason, under a CC BY 4.0 licence.

Which conditions of the June result survived into coverage, across 17 re-verified articles

Which conditions of the AMIE result survived into coverage, 17 articles
Condition the result depends onArticles carrying it
The patients were actors13 of 17
The metric is rating proportions, not accuracy or outcomes9 of 17
Consultations were text chat only8 of 17
The backbone was Gemini 1.5 Flash3 of 17
The paper’s own ablation: newer bare models match or beat AMIE3 of 17
Visits were about two days apart, not weeks1 of 17
Physicians from Canada and India, judged against UK guidelines0 of 17
Source: Ground Truth coding of coverage published 17–30 June 2026, re-fetched and re-checked at origin on 1 September 2026. Full ledger: amie-coverage-ledger.json (CC BY 4.0).
Read before citing
  • The denominator is 17 articles: those that are coverage of the June 2026 disease-management paper, published 17–30 June 2026, and reachable on 1 September 2026 so every coded passage could be re-checked at origin. Three otherwise-eligible articles now return access errors and two had unrecoverable addresses; all are listed in the ledger with their reason.
  • Three further outlets coded by hand against a different rubric — Google's own blog, News-Medical and Digit.in — are listed separately in the ledger and are not in the 17.
  • A condition counts as carried if the article states it anywhere, not only near the headline number.

Zero. That is the condition a referee explicitly challenged, which the authors then wrote into the published paper. We also checked the three articles the counts exclude, coded by hand against a different rubric: Google’s own 17 June blog post, News-Medical and Digit.in. Google’s post names the actors and the specialist raters; News-Medical names the actors, the text chat and the UK guideline design; Digit.in names the actors, the text-only interface and the compressed timelines. None of the three mentions where the physicians practised. The chain is public at every link — referee asks, authors concede, Nature prints it — and it reached zero readers.

Correcting our own first impression, in public: an early sample of six articles suggested a clean escalation ladder from “non-inferior” to “beats.” At 17 that is not what the data show. Fifteen of the 17 carry non-inferiority language somewhere in the text. The press mostly got the verb right. What collapsed was the conditions.

Credit where it is due. The Decoder ran the obsolescence angle in its own headline on 18 June: one result “suggests the tech won’t age well.” The Science Media Centre roundup and Eric Topol’s Ground Truths each carried five of the seven conditions. Ground Truth is not first to the obsolescence point. What is ours is the count, the peer-review file, and the rerun.

In fairness

What would move this rating

The bottom line

Believe this: in a careful simulated study, a Google AI system’s management plans were rated by specialists at least as well as those of 21 experienced physicians, and better on how precisely they were specified and how closely they followed UK guidelines. Its authors said so, and no more.

Don’t believe this: that AMIE has reached an expert standard, or that it beats doctors. The physicians in the June study typed to actors under someone else’s guidelines; the ten physicians in the August study had their cameras off and were scored by raters who could tell the arms apart; “expert-level” is defined nowhere; no study has yet shown a patient better off. The June paper was careful, and its peer-review file was published in full. What failed is downstream of both — and in August, upstream of nothing, because the claim came from the lab itself.

How this was built

Every number above comes from the Nature paper, its supplement, its reporting summary, its peer-review file, the August preprint, or our own runs, read in the source document itself rather than in any summary of it, and each was then re-derived independently. Quotes from the peer-review file carry page numbers. Coverage was coded twice against a fact sheet, and on 1 September 2026 every quoted passage was re-fetched and checked at its origin; articles we could not re-fetch are excluded from the counts and listed in the ledger. The RxQA rerun used the repository’s own prompt at temperature 0, and ships with the scorer, the raw responses with token counts, and the option-shuffle probe. The corresponding authors and Nature received these findings before publication, with time to reply; any response appears here.

Disclosures & provenance

Published
2 Sep 2026
Author
The Ground Truth editor. Editorial standard →
Funding
Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
Rating
Overstated (2/5), rubric v1.0, rated 2 Sep 2026. Full rating card →
Corrections
None to date. Corrections log → · Challenge this analysis