What is Ground Truth?

Before you believe an AI claim in health, check it here.

Every hype claim is a true sentence with the conditions cut out. We put the cut part back, trace it to the primary source, and mark the hidden part in red.

Read the latest investigation → Browse the ratings →

For people deciding what health-AI evidence deserves to be reported, funded, regulated, or deployed.

A hype claim is a true sentence with the conditions cut out. We put the cut part back — in red.

Featured analysis

  • Investigation Benchmarks

    OpenEvidence’s “perfect 100% on MedQA” covers 660 of the test’s 1,273 questions

    OpenEvidence announced on 3 September 2026 that its Darwin model is “the first AI in history to score a perfect 100% on MedQA.” Darwin was scored on 660 of the test’s 1,273 questions: those Google’s physician reviewers most often matched the official answer on before seeing it, after an August review that re-examined only questions some model had got wrong. On the same selection, two other models we ran gain 2.4 and 9.6 points, and Darwin’s lead over Claude Fable 5 is two questions. Darwin’s 660 answers all check out, and on HealthBench Professional, regraded with the benchmark’s own grader, it scores far above physicians’ written answers to the same tasks.

    Read the analysis →

The editorial standard

Credibility you can inspect.

01

Traced to source

Every claim we examine is linked to the primary material — the paper, model card, or registry — so you can check it yourself.

02

Independent by design

We take no money from, and hold no affiliation with, the companies or funders whose claims we scrutinize.

03

Corrected in public

When we get something wrong, we say so on the page, with the date and what changed. Corrections are a feature, not an embarrassment.