The short answer

When you read that a model is “92% accurate,” something had to define the other side of that comparison: the correct answer, the true diagnosis, the real gestational age. That reference is the ground truth. The model’s output is scored against it. So an accuracy number is never a statement about reality on its own — it is a statement about how close the model came to one particular reference, chosen by whoever ran the test.

Change the ground truth and you change the score, without touching the model. That is why, at Ground Truth, the first question about any health-AI claim is not “how accurate?” but “accurate against what — and who decided?

Why it decides everything

Accuracy is only ever as meaningful as the ground truth behind it. A benchmark can be near-perfectly “solved” and still tell you almost nothing about care, if its ground truth is a weak proxy for what matters. Three failure modes recur:

Where ground truth comes from — and how it goes wrong

In health AI, the reference standard usually comes from one of four places, each with a characteristic weakness:

None of these is disqualifying. The point is that the choice is load-bearing, and it is usually made off-stage. A headline reports the score; the ground truth that gave the score its meaning is in the methods section, if it is anywhere.

Ground truth vs. the claim

This is where the term meets our method. A hype claim is a true sentence with its conditions cut out — and the biggest cut condition is almost always the ground truth. “AI matches doctors” is true against a multiple-choice exam, graded by the model’s own maker, measured on no real patient. Put that back and the claim stops being about medicine and starts being about a specific, narrow test. Our recurring analytical move is exactly this: quote the claim, restore the hidden conditions in red, and hand you the question to ask.

How to check a ground truth in one minute

See it in practice

Every Ground Truth investigation is, underneath, an argument about the ground truth a claim rests on:

Or start with the method itself: how to read an “AI beats doctors” claim, how to read a benchmark, and the editorial standard we hold ourselves to.

Disclosures & provenance

Published
18 Jul 2026
Author
The Ground Truth editor. Editorial standard →
Funding
Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
Corrections
None to date. Corrections log → · Challenge this analysis