Explainer
What is “ground truth” in health AI?
In machine learning, ground truth is the reference standard an AI’s output is checked against to decide whether it is right. In health AI it does something quieter and more important: it decides what every accuracy score actually means — and it is where most “as accurate as a doctor” claims come apart.
The short answer
When you read that a model is “92% accurate,” something had to define the other side of that comparison: the correct answer, the true diagnosis, the real gestational age. That reference is the ground truth. The model’s output is scored against it. So an accuracy number is never a statement about reality on its own — it is a statement about how close the model came to one particular reference, chosen by whoever ran the test.
Change the ground truth and you change the score, without touching the model. That is why, at Ground Truth, the first question about any health-AI claim is not “how accurate?” but “accurate against what — and who decided?”
Why it decides everything
Accuracy is only ever as meaningful as the ground truth behind it. A benchmark can be near-perfectly “solved” and still tell you almost nothing about care, if its ground truth is a weak proxy for what matters. Three failure modes recur:
- The reference is a proxy, not the outcome. A model graded against a panel’s label, or against a multiple-choice key, is being measured on agreement with a test — not on whether a patient got better. “Right answer” and “right care” are not the same variable.
- The grader has a stake. When the ground truth is set by the same team, or the same model family, that is being evaluated, the test can be marking its own homework. Independence of the reference is not a detail; it is the difference between a measurement and a press release.
- The reference is narrow. Labels drawn from one site, one population, or selected cases define a ground truth that may not hold anywhere else — so the accuracy number does not travel with the model.
Where ground truth comes from — and how it goes wrong
In health AI, the reference standard usually comes from one of four places, each with a characteristic weakness:
- Expert labels — clinicians annotate the “true” finding. Strong, but only as good as those experts, their agreement, and whether they saw what the model saw.
- Another test — a model is checked against an existing measurement (a biometry scan, a lab value). Convenient, but the reference test has its own error, and “equivalence” is a population average, not patient-by-patient agreement.
- A consensus or answer key — exam-style benchmarks with a single correct option. Cheap to score, easy to contaminate (public questions can leak into training), and a weak stand-in for open-ended clinical work.
- A downstream outcome — did the patient actually do better? The gold standard, and the rarest: most claims never measure it.
None of these is disqualifying. The point is that the choice is load-bearing, and it is usually made off-stage. A headline reports the score; the ground truth that gave the score its meaning is in the methods section, if it is anywhere.
Ground truth vs. the claim
This is where the term meets our method. A hype claim is a true sentence with its conditions cut out — and the biggest cut condition is almost always the ground truth. “AI matches doctors” is true against a multiple-choice exam, graded by the model’s own maker, measured on no real patient. Put that back and the claim stops being about medicine and starts being about a specific, narrow test. Our recurring analytical move is exactly this: quote the claim, restore the hidden conditions in red, and hand you the question to ask.
How to check a ground truth in one minute
- What is the reference? An expert label, another test, an answer key, or a real outcome — and how good is it?
- Who set it? Is the reference independent of the system being tested, or does the developer control both?
- Is it the outcome or a proxy? Agreement with a label is not the same as a patient benefit.
- Where does it come from? One site and one population, or many — and could the test data have leaked into training?
- Was it measured on a patient at all? Most “beats doctors” claims are graded answers, not treated people.
See it in practice
Every Ground Truth investigation is, underneath, an argument about the ground truth a claim rests on:
- Google made Gemma “medical.” We gave it the job. — a medical model that wins on exams (a leaky ground truth) but is matched by a general model on the clinical work that isn’t an exam.
- The scoreboard that can’t be drawn — fifteen “AI beats clinicians” studies whose ground truths don’t compare, so the scores can’t be stacked.
- “Medical superintelligence”: what Microsoft’s 85.5% actually beat — a benchmark graded by the model family being tested.
- “On par with nurses” — a human baseline the company’s own physicians never graded.
- “As accurate as a sonographer” — a genuine result whose ground truth (early-pregnancy dating) is narrower than the headline.
Or start with the method itself: how to read an “AI beats doctors” claim, how to read a benchmark, and the editorial standard we hold ourselves to.
Disclosures & provenance
- Published
- 18 Jul 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Corrections
- None to date. Corrections log → · Challenge this analysis