What is Ground Truth?
Before you believe an AI claim in health, check it here.
Every hype claim is a true sentence with the conditions cut out. We put the cut part back, trace it to the primary source, and mark the hidden part in red.
Read the latest investigation → Browse the ratings →
For people deciding what health-AI evidence deserves to be reported, funded, regulated, or deployed.
A health-AI headline, as published
The cut part, put back
Latest
Featured analysis
-
2.9% or 8%? We recomputed the benchmark behind three “#1” claims
Three clinical-AI companies announced first place from the same Stanford–Harvard clinical-safety preprint in thirteen days. We recomputed the benchmark from the authors' own released per-case data at a pinned commit, using both their statistical code and an independent implementation. The headline safety figure counts case variants, not cases: AMBOSS's cited 2.9% is 8% read as base cases, and the choice of definition moves 21 of 45 models at least three midrank places. On the severe-harm measure the paper's own procedure leaves all four clinical tools in the best set and detects no significant difference among them. AMBOSS does lead the composite score, though its separation from Doximity holds under one test and not another; Doximity leads one 70-case subset by 0.0008; and OpenEvidence's comparison excludes the study-provided assistant, which appears in twice as many records. Scripts, captured outputs and hashes published alongside.
-
“More accurate” than clinicians: what Google’s SymptomAI was actually compared against
Google says clinical experts found SymptomAI’s differentials more accurate than clinicians’. On 517 real conversations with real provider-reported diagnoses, they were — about 73% against about 62%, and the statistics hold.
-
How to read an AI scale-up proposal
So far, clinical AI reliably changes what gets recorded and billed; in the studies reviewed here, we could find no patient-outcome improvement attributable to the AI model itself.
-
How to read a “99% accurate” diagnostic-AI claim
A test can be 99% sensitive, 99% specific, have a 99% negative predictive value, or correctly classify 99% of a selected group.
-
The NHS cancer blood test was “99% accurate.” Which 99%?
-
They want a doctor in every pocket. We handed the pocket a lethal order.
-
AI beat the doctors. So we regraded the doctors.
-
What is “ground truth” in health AI?
-
Google made Gemma “medical.” We gave it the job.
-
“As accurate as a sonographer”: what blind-sweep ultrasound AI actually proved
-
“On par with nurses”: what Hippocratic AI’s headline number actually measured
-
“Medical superintelligence”: what Microsoft’s 85.5% actually beat
-
The scoreboard that can’t be drawn: what 15 “AI-beats-clinician” health studies actually measured
-
How to read an “AI beats doctors” claim
-
“16% fewer errors”: what the OpenAI–Penda Health study actually measured
-
How to read a health-chatbot impact claim
-
How to read an African-language AI benchmark without getting fooled
-
After PEPFAR: autopsy of a health-data empire
Through PEPFAR, the US built and operated the national medical-record systems of a dozen African countries — then vanished in January 2025. The software mostly survived; the donor-paid data workforce did not, and the estate’s one proven benefit — a 28% mortality reduction at Malawi’s HIV clinics — ran through exactly that layer. What replaces it splits three ways: government takeover, sovereignty as a private toll concession (Kenya’s $815M platform contract), and thinly funded drift. AI tools are arriving at precisely the gap the collapse opened.
The editorial standard
Credibility you can inspect.
Traced to source
Every claim we examine is linked to the primary material — the paper, model card, or registry — so you can check it yourself.
Independent by design
We take no money from, and hold no affiliation with, the companies or funders whose claims we scrutinize.
Corrected in public
When we get something wrong, we say so on the page, with the date and what changed. Corrections are a feature, not an embarrassment.
The newsletter
Get the hidden conditions, not the hype.
Every few weeks: one big health-AI claim, traced to its primary source, with the part the headline left out put back in red. That’s the whole email — one-click unsubscribe, no tracking pixels.