What is Ground Truth?
Before you believe an AI claim in health, check it here.
Every hype claim is a true sentence with the conditions cut out. We put the cut part back, trace it to the primary source, and mark the hidden part in red.
Read the latest investigation → Browse the ratings →
For people deciding what health-AI evidence deserves to be reported, funded, regulated, or deployed.
A health-AI headline, as published
The cut part, put back
Latest
Featured analysis
-
OpenEvidence’s “perfect 100% on MedQA” covers 660 of the test’s 1,273 questions
OpenEvidence announced on 3 September 2026 that its Darwin model is “the first AI in history to score a perfect 100% on MedQA.” Darwin was scored on 660 of the test’s 1,273 questions: those Google’s physician reviewers most often matched the official answer on before seeing it, after an August review that re-examined only questions some model had got wrong. On the same selection, two other models we ran gain 2.4 and 9.6 points, and Darwin’s lead over Claude Fable 5 is two questions. Darwin’s 660 answers all check out, and on HealthBench Professional, regraded with the benchmark’s own grader, it scores far above physicians’ written answers to the same tasks.
-
Peer review tempered the claim. The headlines put it back.
Google announced in August 2026 the “first demonstration” of an AI system at “expert-level performance” in real-time video consultations.
-
The AI eye exam can read the retina. Can it save sight?
Retinal AI can identify referable diabetic retinopathy quickly and accurately. Ground Truth traced 29 studies and 15 public claims through the care chain.
-
Health records, then AI: ask what actually moved
Almost every health-technology claim rests on something that moved, and the trick is knowing what.
-
“Predicts risk of more than 300 diseases”: what the Aladynoulli paper reports, and what peer review narrowed
-
2.9% or 8%? We recomputed the benchmark behind three “#1” claims
-
“More accurate” than clinicians: what Google’s SymptomAI was actually compared against
-
How to read an AI scale-up proposal
-
How to read a “99% accurate” diagnostic-AI claim
-
The NHS cancer blood test was “99% accurate.” Which 99%?
-
They want a doctor in every pocket. We handed the pocket a lethal order.
-
AI beat the doctors. So we regraded the doctors.
-
What is “ground truth” in health AI?
-
Google made Gemma “medical.” We gave it the job.
-
“As accurate as a sonographer”: what blind-sweep ultrasound AI actually proved
-
“On par with nurses”: what Hippocratic AI’s headline number actually measured
-
“Medical superintelligence”: what Microsoft’s 85.5% actually beat
-
The scoreboard that can’t be drawn: what 15 “AI-beats-clinician” health studies actually measured
-
How to read an “AI beats doctors” claim
-
“16% fewer errors”: what the OpenAI–Penda Health study actually measured
-
How to read a health-chatbot impact claim
-
How to read an African-language AI benchmark without getting fooled
The editorial standard
Credibility you can inspect.
Traced to source
Every claim we examine is linked to the primary material — the paper, model card, or registry — so you can check it yourself.
Independent by design
We take no money from, and hold no affiliation with, the companies or funders whose claims we scrutinize.
Corrected in public
When we get something wrong, we say so on the page, with the date and what changed. Corrections are a feature, not an embarrassment.
The newsletter
Get the hidden conditions behind the headline.
Every few weeks: one big health-AI claim, traced to its primary source, with the part the headline left out put back in red. That’s the whole email — one-click unsubscribe, no tracking pixels.