Investigation · Benchmarks
“On par with nurses”: what Hippocratic AI’s headline number actually measured
Hippocratic AI, valued at $3.5 billion, sells an “AI nurse.” Its founding paper says the system is “on par with human nurses” — and on average, on its own subjective surveys, it is. But the test used clinicians role-playing patients in simulated calls, a 60-nurse human baseline the company’s physicians never graded, and no real outcomes — and on the safety row, the only advice rated as risking severe harm came from the AI.
In March 2024, the healthcare startup Hippocratic AI published the paper behind its patient-facing “AI nurse.” The centerpiece, stated in the abstract, is that its Polaris system performs “on par with human nurses” on the qualities that carry a patient phone call — medical safety, clinical readiness, patient education, conversation, and bedside manner. That sentence is true to the paper. It is also the load-bearing beam under a company that has since raised $404 million and was last valued at $3.5 billion, and under a marketing story — that the AI matches or beats nurses — that the paper does not support. Read at the resolution it was produced, “on par” means something narrower and stranger than it sounds: near-equal average scores on satisfaction surveys, given by clinicians role-playing patients in simulated calls, against a 60-nurse human baseline the company’s own physicians never graded — a test in which the AI produced the only advice its raters flagged as risking severe harm. This is the number that decides whether a health system lets an algorithm call its discharged patients. It deserves to be read closely.
The claim, and how it grew
The paper — “Polaris: A Safety-focused LLM Constellation Architecture for Healthcare” (arXiv:2403.13313), authored by Hippocratic AI staff — is careful with its own verb. Polaris is “on par with human nurses,” it says, “on aggregate.” Not better; level, on average. The company’s founder and chief executive, Munjal Shah, has been consistent in public that the point is to augment nurses and relieve a staffing shortage, not replace them — a framing worth taking at face value. But between the paper’s hedge and the market’s ear, the verb changed. When Hippocratic announced its partnership with NVIDIA, the coverage carried a blunter promise: the trade outlet New Atlas headlined the deal as agents built to “outperform human nurses,” and reported the company advertising its agents at “less than US$9 per hour,” set against the far higher hourly cost of a registered nurse. The price traces to that one outlet’s reading of Hippocratic’s website; we could find no statement of it in the company’s own materials. By 2026 the company’s own deployment materials were reporting “99.9% clinical safety” and “zero severe harm events.” None of those later phrases is the paper’s. The paper says level; the ecosystem heard victory. This investigation is about the distance between the two — a distance the paper itself, to its credit, gives you the numbers to measure.
What the benchmark actually is
Credit where due: Hippocratic tested a real product on a real task, not a multiple-choice exam. Polaris is a “constellation” — a conversational lead model coordinating specialist support agents for jobs like reading a lab value or checking a medication interaction — and it is aimed squarely at non-diagnostic phone work: post-discharge check-ins, chronic-disease follow-up, medication reminders. To evaluate it, the company ran simulated calls in which more than 1,100 U.S.-licensed nurses and more than 130 physicians posed as patients, then rated the call. After screening out calls shorter than about four minutes, the analysis kept over 3,475 conversations. Two features of that design govern everything downstream. First, the “patients” were clinicians playing a part, and the measures were subjective survey items — how well the caller educated you, whether it covered the critical items, whether it felt safe — not diagnoses made, outcomes changed, or harms averted. Second, the agents are non-diagnostic by design, so the test never asks the thing a nurse’s license is really for: clinical judgment about an undifferentiated patient. Within those walls, the results are genuine. It is the walls that the headline forgets.
“On par,” up close: a two–two split
Aggregate parity hides the shape of the result. The paper breaks the four experience dimensions out, and on them the AI and the nurses split the decision two–two. Polaris was rated higher on bedside manner and on patient education; the human nurses were rated higher on clinical readiness — the single item literally phrased “Was the Nurse/AI as effective as a nurse?” — and on overall conversation quality. “On par” is not the AI edging ahead; it is a near-tie in which each side wins half the categories, on surveys answered by people pretending to be patients.
| Survey dimension (simulated calls) | Human nurses — nurse graders | Polaris — nurse graders | Human nurses — MD graders | Polaris — MD graders |
|---|---|---|---|---|
| Bedside mannerempathy, trust, rapport — Polaris rated higher | 82.67 | 86.32 | – | 84.91 |
| Patient educationPolaris rated higher | 78.48 | 88.24 | – | 90.79 |
| Clinical readiness“as effective as a nurse?” — nurses rated higher | 87.34 | 85.66 | – | 86.7 |
| Conversation qualitynurses rated higher | 85.24 | 84.43 | – | 82.56 |
- These are subjective survey items (for example, “Did the Nurse/AI cover all critical items for this kind of call?”) answered by clinicians posing as patients in simulated phone calls — not clinical outcomes, and not diagnostic accuracy.
- The human-nurse column rests on a subset of 60 nurses in nurse-to-nurse calls — not the 1,100+ nurses who evaluated the AI. The paper reports no confidence intervals or significance tests, so “on par” carries no error bars.
- The entire “MD graders” column for human nurses is empty because, in the authors’ words, they “did not evaluate the U.S. licensed human nurses by U.S. licensed physicians” — only the AI was graded by both nurses and physicians.
- Hippocratic’s agents are explicitly non-diagnostic. “Bedside manner” here means phone and video conversational tone, not physical bedside care.
Two conditions sit under every cell of that table. The scores are satisfaction ratings, not clinical outcomes — a high “patient education” score means a role-playing nurse felt well-informed, which is worth something but is not the same as a patient being safer. And the paper reports no confidence intervals and no significance tests, so “on par” arrives with no error bars: we are told the averages are close, not that they are statistically indistinguishable, nor by how much they could differ if the test were re-run.
The safety row’s tail: who produced the severe-harm advice
The most important row is medical safety, and it cuts in two directions at once — which is exactly why it should be read whole rather than quoted in half. Raters sorted each piece of advice into a harm ladder: nothing incorrect, incorrect but no harm, minor harm, severe harm, death. On the top rung, Polaris looks better: its advice was rated “nothing incorrect” about 97% of the time against the nurses’ 81%. That is the number the “beats nurses on safety” framing is built from, and taken alone it is real — the nurses made more errors, and most of theirs were benign (about 15% “no harm,” 4% “minor harm”).
But keep reading down the ladder. The nurses in this sample produced no advice in the severe-harm or death categories — 0.00% each. Polaris produced some: about 0.04% of its advice was rated as risking severe harm by the nurse graders, and 0.15% by the physician graders. On a rung where the human floor was zero, the AI’s floor was above zero. Both facts are true, and honesty requires both: Polaris errs less often overall, but the only advice anyone flagged as potentially causing severe harm came from the machine.
On a rung where the humans’ severe-harm rate was zero, the AI’s was not — on the company’s own benchmark. For a product sold on safety, a floor above zero is the number to watch, not the one to bury.
The care needed here is real, and it runs against overreach in both directions. This is a small human baseline — a subset of 60 nurses, single-rater, no confidence intervals — so the honest reading is “in this sample the nurses produced no severe-harm advice and the AI produced a little,” not “the AI has a proven higher harm rate.” A larger or differently graded test could move it. But the direction is not nothing, and it is the opposite of the marketing: a system sold on safety produced, on its maker’s own benchmark, the sample’s only advice rated as potentially severe.
Not the same test: the comparison was asymmetric
The parity claim rests on a comparison the paper itself flags as uneven — and the flag is the most quotable line in it. The AI was graded twice, by nurses and by physicians. The human nurses were graded once, by nurses only. In the authors’ own words, they “did not evaluate the U.S. licensed human nurses by U.S. licensed physicians.” So the stricter, higher-status graders — the physicians who rated the AI’s severe-harm share at 0.15% — never marked the humans’ papers at all. And the human side of the whole comparison is thin: not the 1,100-plus nurses of the headline, who were evaluating the AI, but a randomly selected subset of about 60 nurses who took calls from other nurses to form the baseline. A near-tie is being declared between a contestant scored by two panels and a contestant scored by one, on a human sample a fraction the size the headline implies. That does not make the AI worse than a nurse. It makes “on par” a claim the test was not built to settle.
The deployment numbers, and how they are counted
Two years on, the evaluation has been overtaken by scale claims — and the units are slippery. In May 2026 Hippocratic reported reaching “10 million patient calls” at “99.9% clinical safety,” with an average patient rating of 8.95 out of 10. In the same season it also reported “more than 180 million patient interactions.” Those are not the same number: “calls” are voice calls; “interactions” is an undefined, much larger tally — roughly eighteen times bigger — reported alongside the call count as a separate, bigger figure that does most of the headline work. And it climbs fast: the same “interactions” metric stood at about 115 million in November 2025 and 180 million by April 2026, with no definition of what a single “interaction” is.
The safety figures behind those totals are self-graded. “99.9% clinical safety” and “zero severe harm events” are the company’s own metrics, adjudicated by its own paid network of more than 7,500 clinicians — not an independent audit, and not a peer-reviewed result. And “zero severe harm events,” asserted for the 2026 deployment, sits awkwardly beside the company’s own 2024 paper, which measured 0.04–0.15% of advice as risking severe harm. The two rest on different studies and different definitions; they are not a contradiction in the strict sense. But it is the same word — severe harm — carrying the number zero in the marketing and a number above zero in the science.
The evidence base behind a $3.5 billion valuation
Strip out the self-run evaluations and ask what independent, checkable evidence exists. On the record we could find, very little. A search of ClinicalTrials.gov returns no registered trial sponsored by Hippocratic AI. The company’s largest safety study — a real-world evaluation reporting that correct-advice rates rose from about 80% to 99.38% across 307,038 calls — is a non-peer-reviewed preprint, adjudicated by the company’s own 6,234 clinicians, with no control arm. Its two genuinely peer-reviewed papers measure process proxies (test opt-in rates, call metrics), not health outcomes. Marketing materials cite outcome gains such as a reduction in hospital readmissions, but we could not locate a published denominator, control group, study period, or registration for that figure, so we neither confirm nor repeat it. And several of the health systems named as validating customers — Universal Health Services, WellSpan, Cincinnati Children’s — are also investors in the company, named among the participants in its Series C, which does not make their results wrong but does entangle the parties who vouch for them.
Two fair counterweights belong here. The agents really are positioned as non-diagnostic, which keeps them outside the FDA’s device-clearance regime by design and lowers the stakes of any single call; and the largest U.S. nurses’ union has argued, in general terms, that AI in nursing is being deployed ahead of the evidence — a criticism of the category, not a claim about Hippocratic specifically. Both can be true: a tool can be modestly scoped and genuinely useful, and its public numbers can still be running ahead of what has been independently shown.
What would move this rating
Every step is available to the company. Publish the peer-reviewed version of the evaluation. Grade the human nurses the same way the AI was graded — by physicians as well as nurses — and report the human baseline’s size and confidence intervals next to the AI’s. Define “patient interaction” and separate it cleanly from “call.” Register a prospective study on real patients with a control arm and a pre-specified outcome, and let an independent party adjudicate the harm categories rather than the company’s own panel. Above all, close the gap between the paper’s word — “on par” — and the market’s — “outperforms.” The pattern this site watches is a claim that grows in public while its evidence stands still; this is that pattern, in a product already dialing real patients. When a prospective, independently graded trial arrives, we will read it here and re-rate in public, in whichever direction it points. For the reader’s own toolkit, this claim is a natural fit for the seven questions we ask of any “AI beats clinicians” study, and its row belongs beside the others in our scoreboard of what these studies actually measured.
The bottom line
Believe “on par with human nurses” — as a description of near-equal average scores on Hippocratic AI’s own subjective surveys, in simulated calls, where clinicians played the patients and the humans won two of the four experience dimensions. Do not read it, on this evidence, as an AI that matches or beats nurses in the clinic: no real patients were involved, the human comparator was smaller and more leniently graded than the AI, the agents don’t diagnose, and on the safety row the only advice rated as risking severe harm was the machine’s. The engineering is real and the task is legitimate. So is the distance between what the benchmark measured and what the marketing says. When the peer-reviewed trial on real patients arrives, we’ll update this rating — up or down — on the evidence.
Primary sources: Mukherjee et al. (Hippocratic AI), “Polaris: A Safety-focused LLM Constellation Architecture for Healthcare” (arXiv:2403.13313, 2024) · Hippocratic AI, “Polaris 5.0” (2026) · Hippocratic AI × DigitalOcean, “10 Million Patient Calls at 99.9% Clinical Safety” (2026) · Hippocratic AI, RWE-LLM preprint (medRxiv, 2025) · ClinicalTrials.gov · Hippocratic AI, “Safety” (clinician network).
Claim as circulated: New Atlas · Hippocratic AI, Series C announcement (2025).
Ground Truth is independent of, and unaffiliated with, Hippocratic AI, NVIDIA, and their investors. Corrections are made in public — see our policy.
Disclosures & provenance
- Published
- 10 Jul 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Rating
- Needs context (3/5), rubric v1.0, rated 10 Jul 2026. Full rating card →
- Corrections
- None to date. Corrections log → · Challenge this analysis