The registry
Every rating, traceable to source.
Every claim, system, benchmark, and deployment we have rated, newest first. Each entry links to the full rating — the claim quoted verbatim, seven dimensions scored against the primary sources, and the hidden conditions marked in red.
8 ratings · not a leaderboard, ordered by date · machine-readable JSON
-
“The PinPoint AI blood test is 99% accurate at both detecting gynaecological cancers and ruling them out.”
-
“MedGemma models can be adapted to run on mobile hardware and are “small enough to run offline” for health AI — the basis for deploying phone-sized medical AI to frontline care in low-connectivity settings.”
-
The circulating claim that OpenAI's GPT-5.6 outperformed physicians on health evaluations Overstated
“OpenAI's GPT-5.6 outperforms physician responses in health evaluations — 60.5 versus 43.7.”
-
Google MedGemma vs the Gemma base it was built from and a newer general Gemma, on clinical decision support Needs context
“MedGemma: our most capable open models for health AI development.”
-
Blind-sweep ultrasound AI — gestational-age estimation vs. expert fetal biometry Holds up with conditions
“Between 14 and 27 weeks’ gestation, novice users with no prior training in ultrasonography estimated GA as accurately … as credentialed sonographers performing standard biometry.”
-
“Polaris performs on par with human nurses on aggregate across dimensions such as medical safety, clinical readiness, patient education, conversational quality, and bedside manner.”
-
“[The] Microsoft AI Diagnostic Orchestrator (MAI-DxO) correctly diagnoses up to 85% of NEJM case proceedings, a rate more than four times higher than a group of experienced physicians.”
-
OpenAI × Penda Health — AI Consult Needs context
“[Clinicians using the AI Consult copilot had] a 16% relative reduction in diagnostic errors and a 13% reduction in treatment errors compared to those without.”
Composite bands rate the credibility of the claim as stated, not system performance. How we rate →