Building something that behaves like clinical decision support became cheap. Producing evidence that it helps a patient did not. This guide is for the funder on the other side of a queue of near-identical proposals — in a teaching hospital in Boston or a district health office in Kampala — and its central argument is that those are two different questions with two different counterfactuals, so outcome estimates and economics do not transfer automatically between them — though the lessons about measurement, implementation and automation bias do. What follows is what these tools have been shown to do, the two markets you might be funding into, the traps that catch careful readers, thirteen questions for reading a claim, and what the evidence says is worth backing. Every figure links to its primary source.
So far, the clearest measured effect of ambient clinical documentation in billing health systems is on what gets recorded and billed. In the studies reviewed here, we could find no patient-outcome improvement attributable to the AI model itself. Documentation studies report modest time savings on an instrument that may overstate them; the largest controlled cohort found no significant change in after-hours work, while two studies and an industry insight reported increased billing intensity. In the matched cohort reviewed here that measured clinical management alongside documentation, the note and the code diverged without a corresponding change in treatment. One implementation package produced a favourable adjusted mortality estimate while crude mortality was flat, and it did so through staff training, a mandated nurse assessment, required physician communication and ward-level accountability, with a three-variable rule at the centre. A second cut documented errors — not patient outcomes — and only once the same kind of change management was added. Fund the implementation and the measurement, and treat the model as replaceable unless task-specific evidence shows otherwise.
1. What these tools have been shown to do
Documentation is where the evidence is strongest, and it is real. A three-arm randomized trial of 238 UCLA physicians found one ambient scribe cut time in the electronic record's note interface by 9.5% against a concurrent control arm (95% CI −17.2% to −1.8%; p=.02) while a second showed no significant change (−1.7%; p=.66). Two things about that trial travel badly and should be carried with the number. The figure that circulates — 41 seconds saved per note — is the treatment arm's own before-and-after change; the control arm also improved, by 18 seconds, so roughly 23 seconds is the unadjusted difference between those changes (our subtraction on their figures). And the tools were used in only 33.5% and 29.5% of eligible visits, with about 15% of assigned physicians never using them at all, so this is the effect of offering a scribe at roughly a third uptake.
The time instrument has a known blind spot. Epic's Signal metric does not count time spent working inside the scribe application, so reported savings "may represent an overestimation" — a limitation the trial authors note "affects all AI scribe studies reporting Epic's Signal metrics." That includes the largest comparative study reviewed here, an observational cohort rather than a trial: across five academic systems, the 1,809 clinicians who took up a scribe cut electronic-record time by 13.4 minutes and documentation time by 16.0 minutes per eight scheduled patient hours more than the 6,772 who did not, on a difference-in-differences estimate. In that cohort, electronic-record time outside work hours did not change significantly.
Against a human-scribe comparator, the picture is worse. Most scribe studies reviewed here compare against no-scribe care. One compared ambient AI with a human scribe: across 710 emergency department visits, physicians using ambient AI spent more time in the notes section, not less — 4.3 versus 1.8 minutes for adults — contributed 60.1% of the note's characters themselves against 30.8% with a human scribe, and produced notes of similar quality for adults and lower quality for paediatric patients. Note quality against humans is consistent elsewhere: in a Veterans Health Administration evaluation using five standardized cases recorded with actors, notes from 11 AI tools scored lower as a group than notes written by people listening to the same audio, in all five cases, significantly so in three, with the largest gap of 23.5 points on a 50-point instrument in the case with substantial background noise.
The matched cohort reviewed here that looked past documentation at clinical management found a divergence rather than an improvement. In 20,302 matched primary-care annual-visit notes, AI-scribed visits recorded richer neuropsychiatric symptom language and fewer depression diagnoses — 447 visits (9%) carried a depression code, against 587 and 604 (12%) in the human-scribe and contemporaneous unscribed groups and 535 (11%) in visits from before deployment — while new antidepressant prescriptions were essentially unchanged and behavioural-health referrals were, if anything, slightly higher. The composite is driven by coding, not by treatment. That is an important secondary effect, and it is not the same as less care.
For decision support proper, the measured ceiling on patient-facing gains has been low, and the recent evidence has not raised it. A 2020 meta-analysis of 108 studies reporting 122 controlled trials — a pre-LLM literature of reminders, alerts and order sets, with a search closing in August 2019 — found 5.8 percentage points more patients receiving recommended care, and among the 30 trials measuring clinical endpoints, a median 0.3-point change in guideline-target attainment: three patients in a thousand. The registered randomized trial we found that put LLM decision support into live low-resource care and measured it against a patient outcome, across 16 Kenyan facilities, found no significant reduction in 14-day treatment failure (2.2% against 2.0%; adjusted odds ratio 0.77, 95% CI 0.55–1.08), with clinicians fully following the model's advice in 195 of the 1,000 high-severity alerts an expert panel reviewed (19.5%) — alerts the same panel judged definitely safe and appropriate less than half the time.
One 2026 head-to-head found that capability was not confined to specialized products: three frontier models beat two specialized clinical AI products across all three evaluations; on real physician questions the specialized tools scored level with Google's free AI Overview — 3.24 and 3.17 against 3.27. Handing a clinician a model did not reliably add to clinician performance in two trials: in one, physicians with an LLM beat those with conventional resources by two points, not significant, while the model alone beat both, and a later trial found a 6.5% gain on management reasoning with no significant difference between assisted physicians and the model working alone.
2. Two different counterfactuals
In a US health system the counterfactual is a well-resourced clinician who already has an electronic record, may already have a human scribe, already has decision support firing at them, and works inside a billing system. The question is incremental: does this beat what they already have? In a district hospital in a low-income country the counterfactual may be a health worker with a paper protocol, no record system, no billing, an unreliable referral route, and sometimes no clinician at all. The question is categorical, and "better than nothing" is a low bar that sounds compelling and proves little. These are different questions. Outcome estimates and economics do not carry across automatically; the lessons about measurement, implementation and automation bias do.
The high-income question
Here the market is saturated and consolidating. In the Bain/Bessemer/AWS Healthcare AI Adoption Index — a survey of more than 400 senior healthcare leaders published in April 2025 — 30% of provider organisations were deploying AI scribes system-wide and a further 62% were piloting or implementing them. The independent vendors — Abridge, Ambience, Nabla and Suki among them — compete for the same enterprise buyers alongside Microsoft, which reaches them through its Nuance products. Capital and evidence point at different companies: Abridge leads the satisfaction rankings and raised at multi-billion valuations, while Nabla has raised $120 million in total and holds the positive result for one product in the three-arm randomized trial reviewed here. The platform is now a competitor: Epic shipped its own native ambient scribe in February 2026.
The destination is occupied too. The point-of-care answer layer already holds OpenEvidence, UpToDate, Doximity and Elsevier's ClinicalKey AI — which Elsevier built in partnership with OpenEvidence, so one entrant already supplies a rival. Predictive decision support inside the record is the most studied layer and the weakest: a 2026 systematic review and meta-analysis of 22 studies across 34 sites and 2.3 million patients found no Epic model exceeding pooled AUROC 0.79.
And the economics have an answer now, which is the most important development of the past year. Three independent sources point the same way. At UCSF, adopters billed 1.81 more RVUs per week than non-adopters — about $3,044 a year per physician — with the authors themselves calling for work to "ascertain whether increased RVUs reflect more clinical services or accurate coding rather than upcoding." A second health system, Providence St Joseph Health, recorded an immediate rise of 7.40 RVUs per month once clinicians began active use — a within-clinician before-and-after change with no comparison group, not a controlled effect. And PHTI, whose March 2025 assessment was that the financial impact was unclear, now leads its April 2026 industry insight with the opposite: "Provider deployment of AI is increasing billing intensity and inflating medical spending." One multihospital system reported that in its ambulatory and emergency department settings, better documentation and coding produced a 5% rise in Level 5 encounters and a 7% rise in Level 4 encounters for established patients — worth about $1,004 per provider per month. Health plans, PHTI reports, are responding with across-the-board downcoding.
So in a billing system the most consistent signal is a billing one, and a revenue effect is a transfer, not a health gain. The funder's question becomes who pays for it.
The emerging-market question
Almost none of that applies. The structure is inverted: a WHO-authored review retrieved 983 digital health projects across sub-Saharan Africa and analysed 738, finding "an unprecedented level of duplication," with 53.4% of those analysed established. Uganda's health ministry, in its 2023 guidelines for introducing digital health solutions, looks back at a point when there were "over 50 eHealth innovations in Uganda by almost as many donors" — a pattern it calls "pilotitis" — leaving disjointed "information islands," with many projects ceasing to operate once development funding ends. Public infrastructure is central at scale: DHIS2 runs in 80 countries as the government-owned system of record.
One established category has an explicit performance standard. WHO's June 2025 policy statement names six chest X-ray products meeting its tuberculosis screening standards — CAD4TB, qXR, Insight CXR, InferRead DR Chest, DrAid and Genki Edge — and two that did not. Their manufacturers are based in the Netherlands, India, the Republic of Korea, China, Viet Nam and — for Genki Edge — the United States, according to WHO's own table; the two products that failed the standard came from a South Korean and a South African manufacturer. A funder screening this category against a US market map will recognise almost nobody.
Then the arithmetic. Government and donor health spending combined came to US$17.20 per person per year in the median low-income country in 2024 — under US$10 from government alone, against the US$60 the World Bank cites as a minimum benchmark. External aid is both large and shrinking: the World Bank projects off-budget development assistance for health in the median low-income country to fall by about 45% between 2024 and 2030. After the 2025 suspensions, 45% of the 65 WHO country offices that answered the question reported disruption to health management information systems. The randomized-trial base is sparse: of 86 trials of AI health tools run worldwide between 2018 and 2023, four took place in LMICs.
The consequence is structural, and it is the sharpest thing in this guide. The mechanism that makes ambient documentation commercially viable in the United States is billing intensity. That mechanism does not exist in a public health system with no fee-for-service billing. So the American business case does not transfer, and the clinical case that would have to replace it is not established in the evidence reviewed here. A proposal that carries US evidence into an LMIC is carrying the wrong evidence for the wrong question.
3. Traps that catch careful readers
These are properties of measurement rather than of markets, and they apply on both sides of the fork.
The uncontrolled number. A within-arm before-and-after change is not an effect. The 41 seconds above is the treatment arm's own drift; the control arm moved 18. Ask which quantity you are being shown.
The instrument with a blind spot. When the measure of documentation time cannot see time spent inside the tool, the tool can look better than it is. The authors of the randomized trial above say this affects every study using that metric.
The enthusiast denominator. In a large deployment reviewed here — 7,260 physicians across 2,576,627 encounters over 63 weeks — the top third of users accounted for 89% of all activations and 3,447 physicians reached 100 or more encounters. In a survey of 102 adult and family medicine physicians inside that deployment, 8% had never used it. In an emergency department where use was optional, the tool touched 11.2% of eligible encounters. Benefit concentrates; averages conceal it.
The adoption curve that looks like retention. Across the emergency departments of 48 Spanish hospitals, monthly ambient-scribe adoption rose from 7.7% to 57.8% over twelve months, covering 1,032,558 of roughly 2.27 million consultations — an aggregate curve during an ongoing rollout, not a fixed cohort followed from first activation. In the studies reviewed for this guide, we could find no published ambient-scribe cohort that fixes the denominator at first use and reports how many are still using it at month six and month twelve.
The version that moved. Every other question assumes a fixed artifact, and a deployed model is not one: comparing two 2023 releases of the same frontier model on 340 licensing-exam questions, 12.2% of answers changed — one in eight — while the headline score moved only from 86.6% to 82.4%, because the June version corrected old errors and introduced new ones. Epic's sepsis model shows both halves: external validation of the first version returned AUC 0.63, missing 67% of sepsis patients, while a 2026 prospective validation of version two across 227,091 encounters found AUROC 0.82 to 0.92 — materially better, with positive predictive value still 0.13 to 0.26 and the score threshold needed for the same sensitivity ranging from 14 to 37 across four institutions, which is why its authors still advise local validation before deployment.
The transplanted setting. Watson for Oncology's first-choice recommendation matched a Korean tumour board on 41.5% of 65 advanced gastric cancer cases — 87.7% once its second-tier "for consideration" options counted as a match — with the authors attributing most of the discordance to the gap between where it was trained and where it was deployed. NHS England now requires suppliers to show "evidence of real-world benefit … in the NHS care setting proposed."
The pilot that cannot tell products apart. A six-week crossover of two ambient scribes in one emergency department found one tool ahead on burden, note quality and satisfaction — and then found those differences reflected order effects more than the tools themselves, with no overall preference once order was accounted for. In this short head-to-head, the apparent product differences mostly measured which one clinicians touched first.
The safeguard that needs measurement. "A clinician reviews every output" is a design claim with a measured value. In a laboratory experiment, 27 radiologists reading mammograms alongside a mock AI showing pre-drawn heat maps rated 19.8% to 45.5% of cases correctly when the mock-up suggested the wrong category, against 79.7% to 82.3% when it suggested the right one — the study measured susceptibility to a wrong suggestion, not unassisted accuracy, which it never tested. In a multicentre randomized vignette survey of 457 clinicians, a standard model raised diagnostic accuracy by 2.9 points from a 73.0% baseline while a systematically biased one cut it by 11.3, and image-based explanations did not mitigate the harm. And when five platforms were given fourteen identical simulated encounters, a mean 26.3% of key clinical elements were wrong or missing, with omissions about three-quarters of all errors.
4. How to read a claim
Three questions come before the rest, because they decide whether the evidence is worth reading at all. The first two have different right answers on each side of the fork.
- Who is harmed today by the absence of this tool, and what evidence shows this tool prevents that harm? Not what it improves — who has the bad outcome now, and where has the tool been shown to change it.
The state of the evidence: the registered randomized trial of LLM decision support in low-resource primary care reviewed here found no significant reduction in treatment failure; the pooled pre-AI decision-support literature moved guideline attainment by a median 0.3 points; and a 2026 npj Digital Medicine Perspective on scaling these tools reports that ambient-scribe research has focused on clinician efficiency, well-being, time and cost, with no evidence yet of improved clinical or patient-centred outcomes. A proposal that cannot name the harm is offering a capability, not a benefit. - Better than what, exactly? In a high-income system the comparator is the clinician's existing setup, which may already include a human scribe — and against a human scribe, ambient AI cost physicians more time and produced similar-to-worse notes. In a low-resource system the comparator is a trained worker with a working protocol and a referral route, not an empty chair. Ask which comparator the evidence used.
- Is the clinical task hard enough to need AI? Ask what the rule-based version costs and whether anyone has tried it. The largest mortality signal reviewed here came from a three-variable screen — but it arrived wrapped in staff training, mandated nurse assessment, physician communication and ward-level accountability, and crude mortality between arms was flat. Simple and well-implemented has beaten sophisticated and unimplemented more than once.
If a proposal survives those three, the ten below are the ones a real deployment can answer and a demonstration cannot. Most take one email.
- What was the pilot's primary endpoint, and does the scale-up turn on that endpoint? Secondary outcomes need their own prespecification and power; impressions are not endpoints.
What a real answer sounds like: “Time in the note, and it does not tell you anything about diagnosis — here is what we propose to measure for that.” The UCLA trial is a good citizen here: its burnout and safety measures were explicitly secondary. - Which numbers came from the EHR audit log, and which from a survey? Ask for the comparator on the largest one, and whether it is a within-arm change or a change against control.
Why it separates people: in the same randomized trial, the raw difference between two within-arm changes was about 23 seconds, while the adjusted between-group estimate for that product was a 9.5% reduction and the other product did nothing measurable. Survey instruments were warm about both. - Which model version produced these numbers, and who tells you when it changes? Ask for notification or version pinning, a revalidation trigger, and who pays for the revalidation.
The case: Epic's sepsis model went from AUC 0.63 to AUROC 0.82–0.92 between versions. Evidence generated on the first version is not evidence about version two. - What share of eligible clinicians used it, and what share of eligible encounters? Ask for the usage histogram, including never-users and the heaviest third.
The case: at The Permanente Medical Group the top third of users generated 89% of all activations, so a mean saving per clinician is arithmetic performed on a very unequal distribution. In an emergency department where use was optional, the tool touched one encounter in nine. - At month 12, what share of everyone onboarded was still using it? Fix the denominator at first activation; never accept a figure re-based to “active users.”
Bar to clear: in the studies reviewed for this guide, we could find no published ambient-scribe cohort reporting this fixed-denominator retention measure. A new 12-month Spanish cohort reports monthly adoption instead. An applicant who instruments retention before go-live can answer in year two. - Which results came from the setting, acuity and population you are scaling into? Ask specifically for interpreted encounters and the highest-acuity sites.
The case: Watson for Oncology matched a Korean tumour board's recommended treatment in 41.5% of 65 cases — 87.7% if you also count options the board would have considered — with the authors attributing the gap mainly to drugs the Korean national insurer did not reimburse and to practice patterns learned at the American hospital where the system was trained. The model had not changed; the country had. Where care is delivered in a language the model was not built for, ask for measured accuracy, token count, latency and cost in that language. - When the system is wrong, who catches it, and is acceptance measured? Ask whether override rates are logged, and what threshold would switch the tool off.
The case: very experienced radiologists were 82.3% correct when a mock AI suggested the right category and 45.5% when it suggested the wrong one; the study did not measure unassisted accuracy. “A clinician reviews every output” is a claim with a measured value, and the measured value is not reassuring. - Whose budget pays for integration, and what has actually been integrated? Ask for interfaces, analyst hours, and the named owner on the health-system side.
The case: only 30% of completed healthcare-AI proofs of concept in Bain's adoption index reached production, with integration cost among the barriers adopters named. Integration is the buyer's bill, not the vendor's. - What does one encounter cost at full volume, and what is proprietary besides the prompt? Ask for gross margin at projected volume, not at pilot volume.
Why it matters: inference is a cost that recurs on every encounter instead of amortizing, which is why the 70–80%-plus gross margins Bessemer describes as the promise for AI-native health companies remain an aspiration rather than a reported result. Ask for the applicant's own figure at projected volume, and ask what it was at pilot volume. Then ask the opportunity-cost version: what is the cost per patient against the total per-patient budget for this programme, and what would be deprioritised to fund it at scale? - Who computed the savings without the vendor, and what survives if the company fails? Ask which finance office rebuilt the number, and what the contract says about data and service continuity.
Why it matters: the controlled cohort above estimated about $3,044 in additional annual billing per physician. An applicant's projection should be rebuilt by the buyer's finance office, and the contract should specify what happens to data and service if the supplier fails.
5. What is worth funding
The diagnostic half of this guide is easy to mistake for an argument against funding clinical AI. It is not. It is an argument about which parts have been shown to work — and those parts are the same on both sides of the fork, even though what you buy with them differs.
The finding that survives everywhere: the model alone is not the intervention. The largest mortality signal reviewed here came from a stepped-wedge trial that randomized 45 wards across five hospitals and was implemented in 43, covering 60,055 patients in which a three-variable alert was activated alongside staff training, a mandated nurse assessment, a required communication to the covering physician, and feedback to ward leads on how reliably alerts were acknowledged — with crude mortality flat between arms and the benefit appearing after adjustment (adjusted RR 0.85, 95% CI 0.77–0.93). The Nairobi deployment that reduced documented diagnostic errors did so only after active change management: the rate of severe alerts left unresolved sat at 35–40% in both groups while clinicians merely had the tool. Fund the champions, the coaching and the uptake feedback as a named budget line, and report a behavioural uptake metric separately from the clinical endpoint.
Fund measurement that can only be built beforehand. Retention with a fixed denominator, usage histograms including never-users, and a comparator arm cost almost nothing at the start and are impossible to reconstruct later. Make them award conditions rather than reporting requests.
If you are funding a high-income system
Assume a revenue effect and decide in advance who captures it. The billing signal is now the best-replicated finding in the category, and PHTI reports payers already responding with across-the-board downcoding — so a business case built on billing intensity is building on a contested transfer, not on a health gain. Require the comparator that matters: not "before and after," but against whatever the clinician has now, including a human scribe where one exists. Require after-hours burden as a reported endpoint, because that is what clinicians say they want and it is the endpoint the largest study found unmoved. And require local validation before deployment of any predictive model, because performance varies enough across institutions that a threshold from one site is a different tool at another.
If you are funding a low- or middle-income system
Start by naming the mechanism that will pay for it in year four, because the American one is absent. Then fund in the direction the evidence actually supports there. Narrow, high-value case finding with a real denominator: AI-read chest X-ray for tuberculosis is the one category reviewed here with a WHO-referenced performance standard, named qualifying products and priced procurement, and the Nigerian costing study found that combining symptom screening with AI-read imaging detected the most cases at US$773 per bacteriologically-confirmed case, against US$635 for screening on any symptom and US$1,198 for cough of two weeks or more — with the honest note that AI-read imaging on its own was economically dominated, and that symptom screening alone missed 60% of the cases the combination found. Fund integration with the system a ministry already owns rather than a parallel data layer. Fund locally led evaluation, where the funding rails already exist and the gap is measured rather than asserted: four of the 86 randomized AI-health trials identified worldwide between 2018 and 2023 took place in LMICs. And fund the regulatory and procurement capacity that already has a legal mandate, rather than another pilot.
One warning specific to this side of the fork. The capability available to a fully disconnected setting is the least capable tier of the technology, which means the populations with the highest burden are offered the weakest models. That is an argument for narrow, well-specified tasks with a verifiable output — not for a general-purpose assistant in a pocket.
Cut both ways
None of this says the tools do not work. In the three-arm randomized trial, one product reduced time in the note interface against control and survey responses were positive about both, though neither product produced a significant burnout effect. The opposite failure is just as expensive: reading every modest effect as proof of hype, or denying a useful aid to systems desperately short of clinicians because its pilot could not answer a question the studies reviewed here have not yet answered. The absence of outcome evidence is partly because the field is young and these endpoints are hard. That is a reason for precision about what you are buying, not for refusal.
The record, not the patient
On the evidence reviewed here, the measured effects of clinical AI so far are these: modest documentation savings on an instrument that may overstate them; no significant change in after-hours work in the largest controlled cohort; increased billing intensity reported in two studies and an industry insight; a divergence between what notes say and what gets coded in one matched cohort; and no significant patient-outcome benefit in the registered randomized LLM trial reviewed here. The two interventions that came closest produced a favourable adjusted mortality estimate and a lower adjudicated error rate, and they did so through training, mandated communication and accountability, with a cheap rule at the centre and the model nowhere near the middle of the story.
The clearest measured effect so far is on the record. The hoped-for effect is on the patient. A proposal is asking you to fund the second and will be measured on the first.
So ask which counterfactual you are funding against, and demand the evidence that matches it. In a billing system, ask who pays for the revenue effect. Where there is no billing system, ask what replaces it. And in both, fund the implementation and the measurement — because in the interventions reviewed here, implementation was the part associated with the strongest outcome signals.
Frequently asked questions
Do AI scribes improve patient outcomes?
We could find no ambient-scribe study reporting an improved patient health outcome such as mortality, morbidity or readmission, or an improvement in diagnostic accuracy. A 2026 npj Digital Medicine Perspective likewise reports no evidence yet of improved clinical or patient-centred outcomes. The matched cohort we found that examined clinical management alongside documentation found AI-scribed primary care visits recorded richer symptom language but fewer depression diagnoses, with prescriptions and referrals essentially unchanged. That is a coding divergence, not better or worse care, and it is the only such signal we found.
How much documentation time do they actually save?
In a three-arm randomized trial, one product cut time in the note interface by 9.5% against control and another showed no significant change. The largest observational difference-in-differences cohort found 13.4 fewer minutes in the electronic record and 16.0 fewer minutes of documentation time per eight scheduled patient hours. Both studies use an EHR metric that does not capture time spent inside the scribe application, which the trial authors say may overstate savings across every study using it. In the cohort, after-hours record time did not change significantly.
Are they better than a human scribe?
In the 710-visit emergency-department comparison we found, no. Physicians using ambient AI spent more time in the notes section than those with human scribes, wrote more of the note themselves, and produced notes of similar quality for adults and lower quality for paediatric patients.
Why doesn't US evidence transfer automatically to a low-income setting?
Because the counterfactual and the economics both differ. In a US system the comparison is against a clinician who already has an electronic record and possibly a scribe; in a district hospital it may be against a paper protocol. And the mechanism associated with commercial viability in the United States — increased billing intensity — does not exist where there is no fee-for-service billing. There, the business case has to rest on clinical benefit in the target setting; we could find no such evidence in the studies reviewed here.
What has come closest to improving outcomes?
The strongest signals came from implementation packages rather than a model alone. In a stepped-wedge trial, the largest mortality signal reviewed here came from a simple three-variable sepsis screen delivered with staff training, a mandated nurse assessment, required physician communication and ward-level accountability. Even there, crude mortality was flat between arms, 3.2% against 3.1%, and the favourable estimate appeared only after adjustment (adjusted relative risk 0.85). A Kenyan quality-improvement deployment reduced physician-adjudicated diagnostic errors only after active change management: visits ending with an unresolved severe alert ran at 35–40% in both the AI and comparison groups at the start, and fell to 20% in the AI group once coaching and feedback began. Patient-reported outcomes in that study did not differ significantly.
Is a specialized medical model better than a general one?
Not automatically. In a 2026 head-to-head, three frontier general models beat two specialized clinical tools across all three evaluations, and on real physician questions the specialized tools scored level with Google's free AI Overview.
Evaluate a claim yourself
Take the thirteen questions to your own AI
Paste a proposal, concept note, pitch deck summary, or pilot report below. It turns the thirteen questions into a prompt that makes any assistant show its working — separate what was measured from what is being asked for, name the missing denominators, and list what to request before funding. Copy it into ChatGPT, Claude, or Gemini.
You are evaluating a proposal to scale an AI health tool ({{SUBJECT}}), using the "Ground Truth" method (groundtruth.health). I will give you a proposal, concept note, pilot report, or announcement. {{MODE}}
1. Who is harmed today by the absence of this tool, and what evidence shows this tool prevents that harm? Name the person with the bad outcome now, and the study where the tool changed it. The registered LLM decision-support trial reviewed here found no significant reduction in treatment failure.
2. Is it better than a well-supported health worker doing the same task? The comparator is not current state but a trained worker with a working protocol and time to use it. Physicians given a model did not significantly outperform physicians with conventional resources.
3. Is the clinical task hard enough to need AI? Ask what the rule-based version costs and whether anyone tried it. A three-variable screen with a mandated communication step produced a large adjusted mortality signal in a 60,055-patient trial.
4. What was the pilot's primary endpoint, and does the scale-up turn on that endpoint? Secondary outcomes need their own prespecification and power; impressions are not endpoints. Documentation pilots do not test clinical judgement.
5. Which numbers came from the EHR audit log, and which from a survey? Ask for the comparator on the largest one, and whether it is a within-arm change or a change against control.
6. Which model version produced these numbers, and who tells you when it changes? Ask for notification or version pinning, a revalidation trigger, and who pays for revalidation.
7. What share of eligible clinicians used it, and what share of eligible encounters? Ask for the usage histogram, including never-users and the heaviest third. Benefit concentrates; averages conceal it.
8. At month 12, what share of everyone onboarded was still using it? Fix the denominator at first activation; never accept a figure re-based to active users. A fixed-denominator retention curve requires first-activation and subsequent-use logs.
9. Which results came from the setting, acuity and population you are scaling into? Ask specifically for interpreted encounters and the highest-acuity sites. A model encodes the guidelines and payment rules of where it was built.
10. When the system is wrong, who catches it, and is acceptance measured? Ask whether override rates are logged and what threshold would switch the tool off. A wrong suggestion can degrade expert performance measurably.
11. Whose budget pays for integration, and what has actually been integrated? Ask for interfaces, analyst hours, and the named owner on the health-system side. Integration cost lands on the buyer.
12. What does one encounter cost at full volume, and what is proprietary besides the prompt? Ask for gross margin at projected volume, not pilot volume. Inference is cost of goods that recurs on every encounter.
13. Who computed the savings without the vendor, and what survives if the company fails? Ask which finance office rebuilt the number, and what the contract says about data and service continuity.
Prefer primary sources; if a number can only be traced to a press release, a vendor page, or the applicant's own deck, treat it as unverified and say so. Above all, separate what the pilot measured from what the scale-up is asking to be funded, and name the denominators that are missing — who was eligible, who used it, who stopped, and what the comparator was. Do not fill gaps with assumptions — say "not stated" wherever the proposal is silent, and list what you would require before funding.
CLAIM / STUDY TO EVALUATE:
{{CLAIM}}
Ground Truth doesn't run the model for you — deliberately. The method is ours; the judgement stays yours.
Sources
- Documentation time, randomized — 238 physicians, three arms; one product −9.5% against control (95% CI −17.2% to −1.8%; p=.02), the other −1.7% (p=.66); within-arm changes of 41 and 18 seconds; uptake of 33.5% and 29.5% of visits; and the authors' note that Epic's Signal metric excludes time inside the scribe application. Lukac et al., NEJM AI (2025), open version: pmc.ncbi.nlm.nih.gov
- The largest controlled cohort — 1,809 adopters against 6,772 non-adopters across five academic systems; 13.4 and 16.0 fewer minutes per eight scheduled patient hours; record time outside work hours unchanged. Rotenstein et al., JAMA 2026;335(16):1408–1417: pubmed.ncbi.nlm.nih.gov
- Against a human comparator — 710 emergency department visits; 4.3 vs 1.8 minutes in the notes section for adults; 60.1% vs 30.8% of note characters written by the physician; quality similar for adults, lower for paediatric patients. Morey et al., Annals of Emergency Medicine 2026;87:561–568: pubmed.ncbi.nlm.nih.gov. Note quality across 11 tools and five standardized cases recorded with actors, human notes higher in all five: Reddy et al., Annals of Internal Medicine 2026;179:765–772: pubmed.ncbi.nlm.nih.gov
- Clinical management — 20,302 matched primary-care notes; depression codes in 447 visits (9%) with an ambient scribe against 587 and 604 (12%); new antidepressant prescriptions and behavioural-health referrals essentially unchanged. Castro et al., JAMA Psychiatry 2026;83(3):281–286: pubmed.ncbi.nlm.nih.gov
- The decision-support ceiling — 108 studies reporting 122 controlled trials, a pre-LLM literature of reminders, alerts and order sets, search closing August 2019; 5.8 percentage points more patients receiving recommended care, and a median 0.3-point change among the 30 trials reporting clinical endpoints. Kwan et al., BMJ 2020;370:m3216: pubmed.ncbi.nlm.nih.gov
- LLM decision support, randomized — 16 Kenyan facilities; 14-day treatment failure 2.2% vs 2.0%; adjusted odds ratio 0.77 (95% CI 0.55–1.08); full adherence in 19.5% of adjudicated encounters. Agweyu et al., Nature Medicine (2026): nature.com. The Nairobi quality-improvement deployment, and severe alerts left unresolved at 35–40% in both groups until active change management: arxiv.org
- The frontier baseline — three general models beating two specialized clinical tools across all three evaluations, with the clinical tools level with Google's AI Overview on real physician queries. Vishwanath et al., Nature Medicine 2026;32:2405–2409: nature.com. What a model adds to a clinician: Goh et al. 2024 and Goh et al. 2025
- High-income market structure — 92% of surveyed provider organisations deploying, implementing or piloting scribes as of March 2025, and 30% with system-wide deployments, both from the 408-respondent Healthcare AI Adoption Index run by Bain with Bessemer and AWS: Bessemer and Bain. Nabla's $120M total funding; Epic's native ambient scribe; Elsevier's ClinicalKey AI built with OpenEvidence
- Predictive decision support inside the record — 22 studies, 34 sites, more than 2.3 million encounters, no Epic model exceeding pooled AUROC 0.79. Journal of General Internal Medicine (2026): link.springer.com
- Billing intensity, three sources — 1.81 more RVUs per week at UCSF with the authors' own upcoding caveat: Holmgren et al., JAMA Network Open (2026); 7.40 more RVUs per month at a second system: Husa et al., JAMA Network Open 2026;9(5):e2615762; and "provider deployment of AI is increasing billing intensity and inflating medical spending," with a 5% rise in Level 5 and 7% in Level 4 encounters at one multihospital system, in an industry insight drawn from a January 2026 convening: PHTI (13 April 2026)
- Emerging-market structure — 983 projects retrieved and 738 analysed across sub-Saharan Africa, 53.4% established, "an unprecedented level of duplication": Karamagi et al., Journal of Global Health (2022). "Over 50 eHealth innovations … by almost as many donors" and "information islands": Uganda Ministry of Health (2023). DHIS2 as government system of record in 80 countries: dhis2.org
- WHO tuberculosis screening standards — six products meeting the standards and two that did not, with manufacturer countries. "Use of computer-aided detection software for tuberculosis screening: WHO policy statement," 10 June 2025: iris.who.int
- Emerging-market economics — US$17.20 and US$46.60 per capita, the US$60 benchmark, and the aid-versus-domestic comparison: World Bank (19 November 2025). Disruption to health management information systems in more than 40% of WHO country offices: WHO rapid stocktake. Four of 86 randomized AI-health trials in LMICs: Gates Foundation. Cost per bacteriologically-confirmed tuberculosis case: PLOS Digital Health (2025)
- Denominators and durability — 63 weeks, 7,260 physicians, 2,576,627 encounters, top third accounting for 89% of activations, 3,447 at 100 or more encounters, 8% of survey respondents never using it. Tierney et al., NEJM Catalyst 2025;6(5): doi.org/10.1056/cat.25.0040. Optional-use uptake of 11.2% of eligible encounters: Preiksaitis et al.. Monthly adoption rising 7.7% to 57.8% across a 48-hospital network during an ongoing rollout: International Journal of Medical Informatics (2026)
- Version drift and local validation — 12.2% of answers changing between two 2023 releases: Chen, Zaharia & Zou (preprint). Epic sepsis model at AUC 0.63 with 67% of cases missed: Wong et al., JAMA Internal Medicine (2021); version two across 227,091 encounters at AUROC 0.82–0.92, PPV 0.13–0.26, thresholds of 14 to 37, local validation still advised: JAMA Network Open (2026)
- Setting transfer and product comparison — 41.5% concordance with a Korean tumour board across 65 cases: Choi et al.. "Evidence of real-world benefit … in the NHS care setting proposed": NHS England. A six-week crossover in which apparent differences reflected order effects: Webb et al., Applied Clinical Informatics (2026)
- When the system is wrong — 27 radiologists reading with a mock AI, 19.8% to 45.5% correct when the suggestion was wrong against 79.7% to 82.3% when right, with unassisted accuracy never measured: Dratsch et al., Radiology (2023). A systematically biased model cutting accuracy 11.3 points from a 73.0% baseline: Jabbour et al., JAMA (2023). Five platforms given 14 identical simulated encounters, mean 26.3% of key clinical elements wrong or missing: Anderson et al., Mayo Clinic Proceedings: Digital Health (2025)
- The implementation finding — 45 wards randomized and 43 implemented across five hospitals, 60,055 patients; a three-variable screen delivered with training, mandated nurse assessment, physician communication and ward-level feedback; crude mortality flat, adjusted RR 0.85 (95% CI 0.77–0.93). Arabi et al., JAMA 2025;333(9):763–773: pmc.ncbi.nlm.nih.gov
About Ground Truth
Ground Truth is an independent publication that scrutinizes AI claims in health — the benchmarks, accuracy scores, and capability announcements that increasingly decide what gets funded, deployed, and believed. Our aim is not to praise or attack particular products, but to help readers judge the evidence for themselves.
Spotted something we got wrong, or a claim worth taking apart? Corrections and tips are welcome at corrections@groundtruth.health. Read the full editorial standard, independence statement, and corrections policy →
Disclosures & provenance
- Published
- 27 Jul 2026 · last updated 30 Jul 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Corrections
- None to date. Corrections log → · Challenge this analysis