Evidence review
Health records, then AI: ask what actually moved
Almost every health-technology claim rests on something that moved. The trick is knowing what. In African health systems the answer has been remarkably consistent across twenty years and two waves of technology: the documentation moves, the process measures move, and the patient outcome usually does not. The first registered randomized trial of an LLM assistant in African primary care improved the notes and left the patients where it found them — the same shape electronic medical records produced before it, where two decades yielded exactly one rigorous result in which mortality moved. Here is how to tell the difference, and what the evidence actually supports. It is real, and it is not what the sales deck says: the system saved lives by helping clinics find patients who had stopped coming, and the same paper reports no effect on the medicine delivered at the visit. At national scale, meanwhile, the record itself often got worse rather than better. Look wider — at digital decision support and clinical AI across the continent — and about a dozen studies have measured death or a composite containing it. The two trials powered to detect a mortality difference found none, and several point estimates favour the control arm. The first registered randomized trial of an LLM assistant in African primary care landed with that same shape: better documentation, no measured patient benefit.
The test this piece applies, and hands over at the end, is a single question asked five ways: what actually moved? A process measure, a clinical action, or a patient outcome — and who did the work that made it move. The five questions are set out in full, run on a real claim, and available as a prompt you can paste into any assistant. What follows first is the evidence they were built from.
Why this is answerable now
For twenty years, the largest builder of medical-record systems in Africa was not a software company. It was the US AIDS programme. Through PEPFAR, the CDC and USAID paid university partners and contractors to build electronic medical records for HIV care — Malawi’s touchscreen system descended from Baobab Health’s work, KenyaEMR, UgandaEMR, Nigeria’s NMRS, and Zambia’s SmartCare. Most, though not SmartCare, were built on the open-source OpenMRS platform or its data model. In Zambia alone the US put over $300 million into health information systems since 2005. These were built so HIV programmes could count — patients on treatment, retention, viral suppression — because the funder’s reporting demanded it.
That history matters for reading the evidence, because it determined what was measured. Almost everything we know about whether EMRs work in these settings comes from HIV programmes, in a handful of countries, with donor-funded staff doing the data entry. And in January 2025 that arrangement ended abruptly: a US executive order froze foreign assistance, and by August 86% of USAID’s awards had been terminated. Ethiopia lost roughly 10,000 data clerks; South Africa saw more than 15,000 PEPFAR-funded staff furloughed, data capturers among them. The software mostly stayed. The people paid to feed it did not — which, as the evidence below shows, is the part that was doing the work.
The one mortality result, read properly
The strongest study in this literature is a quasi-experimental event study of 106 Malawian HIV clinics, published in the Review of Economics and Statistics in 2025 (open PDF). Adopting the national point-of-care EMR reduced annual deaths by 11% on the conventional estimator and 20% on one robust to staggered rollout, averaged across post-adoption years, reaching 28% by the fifth year — an effect that builds rather than switches on, from about 6% in year one. Effects concentrate in children under ten. The design holds up: no significant pre-trends on mortality, and the result survives four estimators built for exactly this kind of staggered rollout. One caveat belongs here rather than in a footnote, because this piece disqualifies other studies for it: the authors record that in some clinics the EMR arrived alongside a power source and computer-literacy training, so what was measured is that package, not the software by itself.
Read the mechanism, though, because it is the whole story. The authors attribute the gains to efficiency — clinics could see who had lapsed from treatment, trace them, and bring them back. And in the same paper, they report a null: no significant effect on TB testing or treatment, referrals, or CD4 orders, with point estimates often negative, and none on whether an underweight child improved by follow-up. The EMR prompted clinicians and the prompts did not change what they did. What changed was the clinic’s ability to notice an absence and act on it — work performed by people, on foot and by phone.
Two caveats the headline number carries. The 28% is the year-five endpoint, the endpoint of the path rather than its average, identified off the earliest adopters. And the widely quoted $448 per life saved is not a total cost: the paper’s own appendix shows it is a hardware bill of materials — workstations, cabinets, cabling, connectivity — with no line for training, salaries, maintenance or replacement. That is a defensible incremental figure, and the authors argue the point-of-care design deliberately avoided extra staffing. But it is not comparable to the roughly $4,500 per life saved for the most effective charity programmes that the paper sets beside it, which is a full-programme cost. One more thing worth saying plainly: the study is ten months old, with no critique or replication yet — not because it has survived scrutiny — we have seen no independent re-analysis, and the data are proprietary, though the journal requires authors to document an application route.
| Finding | Effect | Design — read this column | Setting | Source |
|---|---|---|---|---|
| National EMR adoption reduced deaths in HIV clinics | 11% (conventional estimator) to 20% (staggered-robust) fewer annual deaths averaged across post-adoption years; 28% by year 5 — the endpoint of the path rather than its average — from ~6% in year 1. ≈5,050 deaths averted by 2019. ≈$448 per life saved is a hardware bill of materials: no training, maintenance, hardware replacement or tracing labour, and no sensitivity analysis varies cost. Not comparable to the ~$4,500-per-life full-programme benchmark set beside it. | Quasi-experimental event study, 106 clinics, not randomized; no never-treated controls (pre-period years k<−6 absorbed into the baseline). Pre-trends null for deaths but untestable for lapses, retention and patients in care — the tracing mechanism the article leans on. Robust to four staggered-rollout estimators. Deaths from clinic records with no vital-registration linkage; better tracing would surface more deaths, so the estimate is conservative on that axis. Data proprietary; we have seen no independent re-analysis, though the journal requires authors to document an application route. | Malawi national rollout — 106 of the 124 EMR sites among 700+ ART clinics, rollout ordered by patient volume; point-of-care model. Partly bundled: the authors note that in some clinics the EMR arrived together with a power source and computer-literacy training, so the estimate is for that package rather than software alone. Effects are similar across clinic types (larger in hospitals), but nothing in the study observes the smallest ~85% of clinics. | source |
| The same rollout did not change the medicine delivered at the visit — a null | No significant effect on TB diagnosis or treatment, referrals, or CD4 orders, with point estimates 'often negative'; none on whether an underweight child improved by follow-up; no significant rise in enrolment. The authors attribute the mortality gain to 'efficiency gains, rather than to changes in the medical care provided at visits.' | Patient-visit-level estimation inside the same event study (Table 2), 106 clinics | Malawi national rollout | source |
| CDSS alerts did change one decision at the visit — and not another | Appropriate action on immunological treatment failure 29.5% → 54.9% (aOR 3.18, 95% CI 1.02–9.87, p=0.05); median time to action 47 → 13 days. Null on appropriate ART initiation (aOR 1.04, 0.41–2.68). | Cluster RCT, 13 clinics, 21,349 ART patients, 2012–14 (NCT01634802); mixed models. Outcome is documented action, and under-recording was 'particularly evident in control sites'; the observed gap fell well short of the 35%→75% the trial was powered for. | Kenya, Siaya County HIV clinics | source |
| Pooled across 25 randomized trials, eHealth moved behaviour; biological outcomes stayed undetected | HIV management behaviours OR 1.21 (95% CI 1.05–1.40). Biological outcomes OR 1.17 (0.89–1.54), p=0.27 — no evidence of effect, not evidence of none. Prevention behaviours OR 1.02 (0.78–1.34). | Systematic review and meta-analysis, 25 RCTs, 10 sub-Saharan African countries. Interventions predominantly SMS and mobile, not EMR or CDSS — contextual evidence, not direct evidence about EMRs. | 10 sub-Saharan African countries | source |
| Most patients traced and interviewed were never lost; a third could not be traced at all | Of 4,560 suspected LTFU, 1,384 (30%) could not be traced. Of 3,176 traced, 952 (30%) had died and 2,224 were alive; of those alive and interviewed, 1,226 (56%) were still on ART. | Retrospective cohort on routine data from an active tracing programme, urban ART clinics, 2006–10 | Lilongwe, Malawi | source |
| The largest test of digital decision support in African primary care put death inside its primary endpoint and found nothing | Co-primary endpoints were (1) a severe complication by day 7 — death, delayed hospitalisation of 24h or more, or hospitalisation without a day-0 referral — and (2) hospitalisation within 24h of a referred consultation. Pulse oximetry plus CDSA, ages 2–59 months: 128 severe-complication events (0.3%), adjusted risk difference 0.1% (95% CI −0.0% to 0.3%); ages 1–59 days: 21 events (0.9%), RD 0.5% (−0.1% to 1.0%). Day-0 hospitalisation with referral, ages 2–59 months: 32 events (0.1%), RD 0.0% (−0.0% to 0.1%). Pulse oximetry alone in Tanzania, ages 2–59 months: adjusted RD 0.1% (−0.1% to 0.3%). In India’s pulse-oximetry arm the same endpoint gave RD 0.1% (0.0% to 0.2%) — the trial’s one interval excluding zero, and it points the wrong way. Null on every other endpoint. The authors report the tools neither increased referral hospitalisation nor decreased deaths or delayed and un-referred hospitalisations, and describe this as contrary to expectations. | Pragmatic, parallel-group, superiority cluster RCT, 172 facilities, 157,677 sick children, 2021–23, published eClinicalMedicine 2025. Tanzania randomised 1:1:1 (control / pulse oximetry / pulse oximetry + CDSA); India randomised 1:1 with NO CDSA arm — it was dropped there pending further adaptation. So the digital-algorithm contrast rests on Tanzania's 24 algorithm facilities against 21 controls; the 157,677 figure is the whole trial and should not be read as the algorithm sample. Intention-to-treat, GEE with facility clusters. NCT04910750, funded by Unitaid, run by Swiss TPH — no vendor stake. Rule-based digital algorithm, not machine learning and not an LLM. The paper reaches the same conclusion independently: it states that evidence on patient outcomes such as hospitalisation and mortality for these tools remains limited and inconclusive. | Tanzania (66 government primary care facilities) and India (106 facilities), sick children 0–59 months | source |
| A digital algorithm cut antibiotic prescribing by 46 points and left the children exactly as they were | Co-primary clinical failure at day 7 (not cured and not improved, or unscheduled hospitalisation): 3.7% (532/14,396) vs 3.8% (543/14,363), adjusted RR 0.97 (95% CI 0.85–1.10) — non-inferior against a 1.3 margin. Co-primary antibiotic prescription: 23.2% vs 70.1%, adjusted difference −46.4 points (−57.6 to −35.2), aRR 0.35 (0.29–0.43). Secondary and exploratory: death by day 7 9 vs 11 (0.1% in both arms), aRR 0.66 (0.24–1.84); non-referred secondary hospitalisation 0.4% in both, aRR 1.14 (0.77–1.69); all hospitalisations by day 7 1.0% (145) vs 0.9% (130), aRR 1.43 (1.00–2.05), p=0.05 — a borderline signal pointing the same way as the Rwandan sibling trial. | Pragmatic, open-label, parallel-group cluster RCT, 20 facilities per arm, 11 months, 44,306 consultations (23,593 intervention / 20,713 control), 28,759 children in the day-7 analysis. Nature Medicine, online 18 Dec 2023, issue 2024;30:76–84. NCT05144763; Fondation Botnar and Swiss Development Cooperation. Day-7 status was ascertained by phone or home visit (day 6–14), so the cure/failure component is caregiver-reported; the hospitalisation component is hard. Death and hospitalisation sit under secondary and exploratory outcomes, not co-primary. Rule-based tablet algorithm, not machine learning. Read this as a SAFETY result, not a benefit result: children were not made better, they were made no worse on two-thirds fewer antibiotics. | Tanzania, Mbeya and Morogoro regions, 40 primary care facilities, children 1 day to 15 years | source |
| The same algorithm without randomisation: prescribing fell just as far, and hospitalisations doubled | The sole primary outcome was antibiotic prescription — a process measure: 70.5% to 24.5%, −46.0 points (95% CI −52.5 to −39.5) per protocol (ITT 80.0% vs 43.3%, −36.7). Secondary day-7 outcomes: clinical failure aRR 1.07 (0.97–1.18), meeting the pre-specified non-inferiority margin of 1.3; but referrals aRR 2.11 (1.20–3.70), primary hospitalisations aRR 2.02 (1.21–3.38), secondary hospitalisations aRR 1.93 (1.17–3.16), and severe outcomes 0.66% (43/6,520) vs 0.36% (25/6,932), aRR 1.88 (1.11–3.15), p=0.018. The before–after comparison shows none of it (severe outcomes aRR 1.64, 0.94–2.85; clinical failure 1.00, 0.92–1.10), which is why the authors read the excess as baseline case-mix rather than harm. | Pragmatic, open-label cluster NON-RANDOMISED controlled trial — allocation by geography, 16 health centres per group, 59,921 consultations enrolled and 47,822 analysed, Dec 2021–Apr 2023, PLOS Medicine 26 Feb 2026. NCT05108831; Fondation Botnar and the Swiss Agency for Development and Cooperation. That the safety picture blurs the moment randomisation is dropped IS the row: the authors' confounding explanation is plausible and untestable. Day-7 loss to follow-up was 30% overall and differential (34% control vs 26% intervention), which weakens the comparison further. Do not report the 1.88 as established harm; do not report the trial as a flat null either. | Rwanda, 32 health centres, paediatric outpatients | source |
| A randomized trial of mobile decision support had neonatal mortality as its primary endpoint, and the estimate moved the wrong way | Institutional neonatal mortality, the pre-registered primary outcome: adjusted OR 2.09 (95% CI 1.00–4.38, p=0.051), intention-to-treat. Intervention districts went from 4.5 to 6.4 deaths per 1,000 deliveries; control districts from 3.9 to 4.3. 198 neonatal deaths in intervention districts against 150 in control, 348 in total. The authors conclude that giving frontline health workers access to the system did not improve institutional neonatal mortality. | Cluster RCT, 16 districts randomised 8/8, 18-month intervention (Aug 2015–Jan 2017), 65,831 institutional deliveries, mortality extracted from DHIS-2. NCT02468310 and PACTR20151200109073; Netherlands Foundation for Scientific Research (WOTRO) and Utrecht University. Read as a firm null with an uncomfortable point estimate, NOT as evidence of harm — the confidence interval touches 1.00. Two things blunt it further: the intervention was a four-part mobile package (free voice calls, SMS, data, and USSD protocol retrieval) of which only about 2% of phone use was protocol retrieval, so the decision support was barely delivered; and the authors flag weak birth-and-death registration, patient flow across cluster boundaries and unmeasured confounding. Not an EMR and not AI. | Ghana, Eastern Region, 16 districts | source |
| A trial designed and powered on neonatal deaths averted none | Neonatal mortality, the pre-registered primary outcome: 18.8 per 1,000 live births with two-way SMS against 15.2 per 1,000 with usual care; RR 1.25 (95% CI 0.81–1.91), P=0.31 — the point estimate favours usual care. 83 neonatal deaths. No intervention-related adverse events. The authors note most neonates died before facility discharge, which caps what any outpatient messaging tool could plausibly change. | Parallel, unblinded, individually randomised 1:1 trial, intention-to-treat, six health facilities; 5,020 randomised Sept 2020–June 2022, 136 excluded for incomplete mortality data leaving 2,442 per arm. NCT04598165, Nature Medicine 2025. The intervention is automated maternal and newborn SMS from 28–36 weeks through 6 weeks postpartum with free-text access to a nurse — not an EMR, not a CDSS, not AI. Include it because it is the cleanest mortality-powered digital-health trial in the region and it is null. | Kenya, six health facilities, pregnancy through six weeks postpartum | source |
| Decision support layered on a digital record in HIV care: non-inferior, not superior, and no mortality difference | Primary endpoint — engagement in care WITH documented viral suppression (<50 copies/mL) at 24 months, a composite of retention and virology: 77.9% (2,649) vs 74.3% (1,759), adjusted OR 1.18 (95% CI 0.95–1.46) against a non-inferiority margin of 0.8. Non-inferior; superiority not demonstrated. Safety endpoints: all-cause mortality 80 (2.4%) vs 53 (2.2%), adjusted HR 1.10 (0.78–1.58); incident tuberculosis 15 (0.4%) vs 14 (0.6%), aOR 0.70 (0.30–1.63). The only thing that moved was disengagement from care, 156 (4.6%) vs 167 (7.1%), aOR 0.67 (0.48–0.93) — retention again, the same mechanism as Malawi. | Pragmatic, open-label, parallel-group cluster-randomised non-inferiority trial, 18 rural nurse-led clinics, enrolment Oct 2020–Mar 2022, 24-month follow-up; mITT 5,770 (3,401 / 2,369) of 5,809 enrolled. eClinicalMedicine 2026. NCT04527874; Moritz Straus-Foundation and Swiss National Science Foundation. The control arm already had digital documentation, so the contrast is decision support plus messaging ON TOP of an existing electronic record — the closest analogue in the literature to the question this article asks. The authors flag that mid-trial changes to national ART guidelines may have narrowed the gap. | Lesotho, 18 rural nurse-led HIV clinics, Butha-Buthe and Mokhotlong | source |
| The largest digital-adherence trial put death in its primary composite and found nothing in either African country | Primary composite poor end-of-treatment outcome — documented treatment failure, loss to follow-up of two or more consecutive months, switch to a multidrug-resistant regimen, or death: South Africa adjusted OR 1.19 (95% CI 0.88–1.60, p=0.25); Tanzania 1.49 (0.99–2.23, p=0.056), the point estimate favouring standard care; Philippines 1.13 (0.72–1.78, p=0.59); Ukraine adjusted RR 1.15 (0.83–1.59, p=0.38). The authors state the technologies did not reduce poor treatment outcomes in any of the four countries. Two social-harm incidents were reported from inadvertent disclosure of TB status by the pillbox, both leading to withdrawal. | Four independent pragmatic cluster-randomised trials, 220 clusters randomised 1:1, intervention clusters further randomised to a smart pillbox or SMS medication labels (except Ukraine); 25,606 enrolled June 2021–July 2022, 23,483 in ITT. Two social-harm incidents were reported in the pillbox arm, from inadvertent disclosure of treatment status, and both participants withdrew. The Lancet 2025; ISRCTN17706019; Unitaid-funded. Two riders: the composite is diluted by loss to follow-up, a programmatic classification rather than a clinical event; and smart pillboxes and SMS labels are neither AI nor an EMR — they belong in the file as digital-health context, clearly labelled. | South Africa and Tanzania (plus Philippines and Ukraine), primary care TB treatment facilities | source |
| Same technology, third African country, twelve months of follow-up including relapse — still nothing | Primary composite at 12 months from treatment initiation (death, loss to follow-up, treatment failure, switch to drug-resistant treatment, or TB recurrence): smart pillbox adjusted OR 1.04 (95% CI 0.74–1.45); SMS medication labels aOR 1.14 (0.83–1.61). The authors state the interventions showed no reduction in unfavourable outcomes. Results were consistent in complete-case and per-protocol analyses; the labels arm showed weak evidence of reduced loss to follow-up only. | Three-arm pragmatic cluster-randomised controlled trial, 78 facilities randomised 1:1:1, multiple imputation for the composite, 3,858 in ITT of 3,885 enrolled. Lancet Digital Health, online 29 Sept 2025, DOI 10.1016/j.landig.2025.100895. PACTR202008776694999; Unitaid. Adjusted risk differences +0.96 points (95% CI −1.19 to 3.11) for the pillbox and +0.42 points (−1.75 to 2.59) for the labels; both intervals span zero. Adherence technology, not AI or an EMR. | Ethiopia, 78 facilities treating adults with drug-sensitive TB | source |
| Adherence rose thirty points. Whether patients were cured or died did not move. | The primary outcome was adherence of 80% or more measured by medication-monitor openings — a process measure: 81.0% intervention vs 50.8% standard of care, adjusted RR 1.51 (95% CI 1.36–1.66). The secondary composites containing death did not move: unfavourable outcome at 18 months 17.1% (216/1,096) vs 22.3% (216/974), aRR 0.78 (0.53–1.16), p=0.21; end-of-treatment outcome aRR 0.82 (0.56–1.22), p=0.31. Deaths made up 72 of the 216 unfavourable outcomes in the intervention arm and 57 of 216 under standard care. | Cluster-randomised trial, 18 primary health clinics across three provinces, May 2019–Feb 2022; 2,727 enrolled, 2,584 with adherence data, 2,070 in the 18-month complete-case analysis. eClinicalMedicine 2024. PACTR20190268115772; Gates Foundation, Stop TB Partnership and the South African Medical Research Council. Control participants also carried monitors, silently, so the contrast is the differentiated-care response rather than the device. This is the cleanest within-trial demonstration of the process–outcome gap in the entire file: the process endpoint moved thirty points inside the same trial in which the clinical composite did not move at all. | South Africa, 18 primary health clinics, drug-susceptible TB | source |
| The largest effect in this table came from the arm with a trained human support team attached | Primary composite of died, failed treatment, or lost to follow-up, against a control rate of 12.4%: platform plus a trained human support team −2.6 percentage points (95% CI 1.0–4.2); platform alone −1.9 points (0.3–3.4); daily SMS reminders alone −1.9 points (−0.1 to 4.0), crossing zero and null. The authors report the reductions came mostly from loss to follow-up; mortality is not broken out anywhere. An unannounced urine-isoniazid substudy in 731 participants showed 7.5 fewer points of non-adherence in the Keheala arm (2.6–12.5). | Open-label four-arm RCT with unequal 4:3:12:12 central allocation across 902 clinics; 16,753 randomised Apr 2018–Dec 2019, 14,962 in mITT. NCT04119375; USAID Development Innovation Ventures. Three things must appear wherever the 2.6 points appear: it is a medRxiv preprint posted 14 Feb 2026, six years after randomisation ended and not peer-reviewed; the manufacturer ran the evaluation — Keheala's founder and chief executive is an author, alongside its project manager and six further authors from its study team; and the composite is dominated by loss to follow-up, a programmatic classification, not by death. Include it anyway: it is the counterexample a critic would raise, and the arm that won is the one carrying a paid human support layer. | Kenya, 902 TB treatment clinics | source |
| Deep-learning chest X-ray reading nearly doubled TB treatment starts and left 56-day mortality where it was | Secondary outcome, 56-day all-cause mortality: 26.1% (54/207) with enhanced diagnostics vs 25.0% (52/208) with usual care, HR 1.05 (95% CI 0.72–1.53). The primary outcome was TB treatment initiation before death or discharge — a process measure: 22.2% (46/207) vs 11.5% (24/208), RR 1.92 (1.20–3.08). The authors attribute the treatment-initiation gain mainly to urine Determine-LAM rather than to the computer-aided X-ray, and note that inpatient mortality for adults with HIV remains very high. | Cluster-randomised trial with admission days as clusters, 4:4:1 allocation across 207 clusters, 415 participants in mITT, single hospital. Online May 2024, print Clin Infect Dis 2025;80(5):1143–51. NCT04545164; Wellcome 203905/Z/16/Z. The AI is confounded by design: CAD4TB v6 was bundled with two urine LAM tests inside one arm, so this is NOT a clean test of the algorithm. It is, however, a genuine null mortality result from a deployed deep-learning tool in African care, and both readings hold. | Malawi, Zomba Central Hospital, hospitalised adults living with HIV | source |
| Computer-aided X-ray got people onto TB treatment ten days sooner without a measurable change in deaths — five years before the current wave | The primary outcome was time to TB treatment, a timeliness measure rather than a health one: median 11 days under standard care against 1 day with HIV-TB screening, HR 2.86 (95% CI 1.04–7.87), p=0.04. The patient outcomes were secondary and null. The paper states there were no significant differences between any pair of arms in all-cause mortality by day 56, and reports no per-arm death counts or effect estimate — so only the null statement is quotable, never a number. Undiagnosed or untreated bacteriologically confirmed TB at day 56: 0.5% (2/382) standard care, 1.0% (4/414) HIV screening, 0.5% (2/410) HIV-TB screening, RR 0.93 (0.13–6.58) — a confidence interval wide enough to be uninformative. | Open three-arm pragmatic randomised trial, 1:1:1, 1,462 adults (473 / 492 / 497). CAD4TB v5 with sputum Xpert triggered by a high score. PLOS Medicine 2021; NCT03519425; Wellcome Trust. Falls outside a 2023–26 window and should be cited explicitly as pre-LLM-era precedent: the pattern of a faster process with no demonstrated health gain is not new, and the patient outcomes here were secondary and badly underpowered. | Malawi, adults aged 18 and over with cough at acute primary care services | source |
| A digital tool built explicitly to improve newborn survival did not detectably improve it | Neonatal mortality among admitted neonates after implementation: RR 0.877 (95% CI 0.541–1.423, p=0.596) — null. Low-birth-weight subgroup (1.5–2.5 kg): RR 0.356 (0.127–1.002, p=0.051), pre-implementation 18.25% against 7.60% post. The economic evaluation reports $28.44 per healthy life year gained within the study period, falling to $6.35 at scale, against a Zimbabwe cost-effectiveness threshold range of $17–855. | SINGLE-GROUP interrupted time series with economic evaluation at ONE hospital, no concurrent control — 2,879 neonatal admissions and 425 deaths, March 2020 to October 2023, implementation December 2020. BMJ Global Health, 9 Jan 2026. Vulnerable to secular trends and not comparable in strength to the cluster-randomised rows above. Cite as a null, never as evidence of benefit; the low-birth-weight result is a subgroup finding at p=0.051 in an uncontrolled design, which is the weakest possible basis for reading it as suggestive. | Zimbabwe, Chinhoyi Provincial Hospital neonatal unit, Neotree digital data capture plus decision support | source |
| An electronic registry with decision support improved every process measure it targeted and moved no health outcome (conducted outside Africa) | Primary health outcome, a composite of conditions at delivery (moderate or severe anaemia, severe hypertension, large-for-gestational-age, undetected small-for-gestational-age or malpresentation): 21.7% (700/3,219) vs 21.9% (688/3,148), adjusted OR 0.99 (95% CI 0.87–1.12). Secondary: stillbirth 7 vs 6 per 1,000, aOR 1.07 (0.57–2.00); moderate or severe anaemia 1.1% vs 1.4%, aOR 0.82 (0.51–1.31); severe hypertension 0.3% vs 0.5%, aOR 0.61 (0.27–1.36). Every primary process outcome moved: guideline-adherent anaemia screening and management 44.3% vs 28.9%, aOR 1.88 (1.52–2.32); gestational diabetes 50.7% vs 39.7%, aOR 1.45 (1.14–1.83); hypertension 96.6% vs 94.7%, aOR 1.62 (1.29–2.05). | Pragmatic cluster-randomised superiority trial, 133 government clinics forming 120 clusters (60 vs 59 analysed), 6,367 pregnant women; the trial ran BOTH process and health outcomes as primary, and stillbirth is a secondary, not part of the primary composite. Lancet Digital Health, 25 Jan 2022. ISRCTN18008445; European Research Council and Research Council of Norway. Conducted in the West Bank rather than Africa, and included here as the closest EMR-plus-decision-support analogue in a comparable health system. The authors note only 9.4% of women attended the full antenatal schedule, which caps what any in-visit tool could have achieved. Process denominators are visits, not women. | West Bank, Palestine — outside Africa; the closest EMR-plus-decision-support analogue in any lower-middle-income setting | source |
| A text-message and voucher link to the health system was followed by fewer perinatal deaths — the one randomized signal pointing this way | Perinatal mortality 19 per 1,000 births in intervention clusters against 36 per 1,000 in control clusters, OR 0.50 (95% CI 0.27–0.93). This was a SECONDARY outcome; the trial’s primary outcomes were antenatal-care attendance and skilled delivery attendance. Not powered for mortality. | Pragmatic cluster-randomized controlled trial, 24 primary health care facilities across six districts, 2,550 pregnant women (1,311 intervention / 1,239 control) followed to 42 days after delivery. Not an EMR, decision-support or AI tool — a messaging and free-call voucher intervention. | Zanzibar, Tanzania — antenatal care at primary health facilities | source |
At national scale, the record often got worse
The case for EMRs usually rests on data quality, and at flagship sites that case is strong: when Uganda’s Infectious Diseases Institute moved from clerk transcription to provider-entered data, recorded discrepancies on four HIV data elements fell from as high as 94% to under 13%. But that is an uncontrolled before-and-after at one elite institute across four years, during which the clinic also added a real-time quality-assurance programme and dropped its standardised forms. It does not describe what happens at scale, and at scale the picture inverts.
In Rwanda, a comparison of 3,467 records across 50 facilities found the EHR less complete than the paper it replaced: viral-load results present in 76.4% of electronic records against 94.7% on paper, viral-load dates matching in under a third, and only two of fifty facilities reaching acceptable concordance. In Kenya, an assessment across 53 KenyaEMR facilities found the system agreeing with its own paper source on 11.9 of 20 data elements, with 31% of records carrying at least one missing value among nine required fields — and, in the 27 facilities reassessed after a dedicated data-quality intervention, 13.6 of 20. That is the ceiling a deliberate fix achieved.
Worse, the outcome that funders care about most is the one the records get wrong. A collaborative analysis of 505,634 patients across 57 African cohorts moved five-year retention from 52.1% to 66.6% once routine records were corrected using tracing data — on the assumption, imported from a separate meta-analysis of tracing studies, that 20.8% of those recorded as lost to follow-up had died and 35.9% had silently transferred. Those weights carry wide intervals (mortality 11.3–35.1%), so the correction follows from the assumption rather than testing it. A South African study of eight primary-care facilities using the national HIV database found 36% of patient outcomes misclassified. The database does not know what happened to the patient; the person who goes and looks does.
And the third pillar of the usual case — that EMRs slash reporting burden — turns out to have no measured evidence behind it at all. The often-cited figure, that Malawi’s quarterly cohort reports once took up to five days of clinic-closing manual counting and became a button-press, comes from a programme report written by the system’s own developers. The five days is an uncited, hedged upper bound that was never timed, with no before-and-after measurement. The same page records that the generated reports came out wrong, from “incorrect software logic,” and took several iterations to fix. We could not find a study that measures reporting time in these settings.
| Finding | Effect | Design — read this column | Setting | Source |
|---|---|---|---|---|
| Provider entry cut recorded errors against clerk transcription — at one flagship institute | Recorded discrepancies fell from 51.9–94.1% to 0.9–12.5% on four HIV data elements; not attributable to provider entry alone. | Uncontrolled before/after at one institute across a four-year gap. 2007: retrospective review, 100 patients / 2,382 visits. 2011: prospective, 10,920 patients / 34,957 visits. Concurrent changes the authors concede: a real-time QA programme from 2008, standardised forms dropped in March 2011 mid-window, and four years of accumulated staff experience. After the switch the free-text note used as reference standard and the database entry originate with the same clinician, though extraction was by independent reviewers. Repeated visits treated as independent in significance testing. | Uganda, Infectious Diseases Institute | source |
| At national scale the EHR was less complete than the paper it replaced | Viral-load result completeness 76.4% in the EHR vs 94.7% on paper; viral-load dates matched in 32.4% of records; regimens matched 82.1%; mean concordance 10.2 of 15 variables; only 2 of 50 facilities met the 85% threshold. | Cross-sectional paper-vs-EHR record comparison, 50 facilities, 3,467 records, 194,152 data items | Rwanda, OpenMRS HIV care | source |
| A national EMR agreed poorly with its own paper source, and a dedicated fix barely moved it | Concordance 11.9 of 20 data elements with 31% of mandatory elements missing at baseline; 13.6 of 20 and 13% missing after a data-quality intervention (missingness risk ratio 0.43, 95% CI 0.32–0.58). | Prospective before/after routine data-quality assessments, 27 facilities, 2,369 then 2,355 records, 2014–15 | Kenya, KenyaEMR | source |
| A national HIV database misclassified a third of patient outcomes | 36% of outcomes wrong: 49.6% recorded as LTFU against 12.2% truly LTFU after tracing; 40% of deaths and 43% of transfers missed. Undocumented transfers were the largest contributor. | Record review plus tracing of 1,074 patients classified LTFU, 8 public primary-care facilities, 2014–17 | South Africa, rural Mpumalanga, national TIER.Net | source |
| Routine records misclassify who is lost, who transferred and who died | Crude 5-year outcomes 52.1% retained / 41.8% LTFU or stopped / 6.0% died, corrected to 66.6% / 18.8% / 14.7%. Among patients recorded as LTFU, 20.8% had in fact died and 35.9% had silently transferred. | Collaborative analysis, 505,634 patients initiating ART 2009–14 across 57 cohorts; inverse-probability weighting from a meta-analysis of 32 tracing studies covering 20,365 traced patients | Central, East, Southern and West Africa | source |
| Most authorized accounts on the national EMR are inactive; fully paperless sites are rare | Mean 18.1% of authorized accounts active per month (SD 13.1; range 7.3–46.8%) — the authors attribute much of this to dormant accounts. Entry modes: 9.4% fully paperless, 52.6% hybrid, 38.0% fully retrospective. | Census attempt across all 376 KenyaEMR facilities; 312 authorised participation and 213 of those (68.3%) returned server-log data, 19 counties; usage extracted from server logs, not self-reported. Data span 2012–19, collected 2020. Facilities where implementation had failed were excluded, so this overstates use. | Kenya, KenyaEMR | source |
| EMR implementation roughly doubled patient time in clinic | Total patient clinic time 37→81 and 56→106 minutes per visit, about three-quarters of the increase being waiting. Countervailing: nurses' patient-care time fell 60–80%, clerks' fell, and outpatient visits rose 85% (404→749/month). | Before/after time-motion studies, three rural health centres, 2008–11. The authors state concurrent growth in visits and staffing confounds the time results, and that they made no attempt to assess quality of care. | Kenya, AMPATH primary care | source |
| End users at a private tertiary hospital preferred the EHR to paper | 89.9% said the EHR made work better than paper; 82.3% disagreed that it added work; only 55.4% found downtime acceptable. | Cross-sectional staff survey, 471 of 548 respondents (85.9%). Perception, not an outcome; one private tertiary hospital. | Kenya, Aga Khan University Hospital Nairobi | source |
| Implementers report cohort reporting went from days of counting to on demand — and that the reports came out wrong | 'Up to 5 days' of clinic-closing manual counting replaced by generated reports — an uncited, hedged upper bound, never timed, with no before/after measurement. From the same report: data-entry errors at the point of care, and 'errors in reports due to incorrect software logic' that 'took several iterations to resolve.' | PLoS Medicine 'Health in Action' programme report — not a study — co-authored by the system's developers (Baobab Health Trust), competing interests declared as none. 6 deployment sites, 42,834 registered patients as of Dec 2009. | Malawi, touchscreen point-of-care EMR in ART clinics | source |
| A co-designed public-sector electronic register was abandoned after going live | Roughly a year live across three facilities, then stalled. Named causes: a change of developers, withdrawal of the NGO funder, provincial capacity gaps, and mistrust stemming from corruption and abuse of the tender system. | Qualitative CFIR case study, 38 interviewees across managers, implementers and end users. An electronic version of a paper primary-care service register, not a full clinical EMR. | Ekurhuleni district, Gauteng, South Africa | source |
| Alerts plus a tracing workforce cut documented loss to follow-up | 45.6% → 32.8% documented LTFU (aOR 0.70, 95% CI 0.65–0.77). Relinked to care: 23.3% control vs 30.6% intervention. | Secondary analysis of the same 13-clinic trial, 5,901 patients. Both arms had the EHR — the contrast is alert-on vs alert-off, actioned by community social workers by phone and home visit. Outcome is documented LTFU, which the alerts explicitly told staff to revise. 90-day LTFU definition (~10% misclassification vs 7.7% at 180 days). About a third of records excluded for missing follow-up data. GEE named; no ICC or correlation structure reported. | Kenya, Siaya County HIV clinics | source |
Deployed is not used
Underneath all of this sits a usage problem that most coverage statistics hide. A census attempt across Kenya’s 376 KenyaEMR facilities, which returned usable server logs from 213 of them, found a mean of 18.1% of authorized accounts active per month — the authors attribute much of that to dormant accounts — and that only 9.4% of facilities were fully paperless, with 52.6% hybrid and 38.0% entering everything retrospectively. In Uganda, 94.3% of upper-tier public facilities run paper and electronic systems side by side, with an average of 4.81 systems each and only 1.9% fully digital.
The cost side is thinner than it should be but points one way. Time-motion studies at three rural Kenyan health centres found total patient time in clinic roughly doubled after EMR implementation — 37 to 81 minutes at one site, 56 to 106 at another — with about three-quarters of the increase spent waiting, though the authors caution that concurrent growth in visits and staffing confounds the result. What fell was not clerical burden but patient-facing time: nurses’ and clerks’ patient-care time dropped after implementation, by around 80% at one site.
AI meets the same test — and produces the same shape
Into this arrives clinical AI, and the argument for it is strong on paper: if the value ran through documentation and follow-up rather than through changing clinical decisions, then tools that write, check and chase are aimed at the right target. The evidence base is now good enough to test that. For LLM-based tools it comes overwhelmingly from one deployment, the Penda Health primary-care chain in Kenya. For rule-based digital decision support — the older, unglamorous kind that encodes a guideline as a flowchart — the African evidence base is far larger and almost entirely null, which is the context every LLM claim should be read against.
The first result to circulate widely was the 2025 study of “AI Consult”, an LLM safety net embedded in Penda’s EHR, reporting 16% fewer diagnostic errors and 13% fewer treatment errors — a claim we examined when it first appeared. Three things about it are easy to get wrong. It is often described as observational. It was not: within each of 15 clinics, clinicians were randomly assigned to have the tool or not, 57 with and 49 without — though the paper calls itself a pragmatic cluster-assigned study, and it was never registered. The errors are physician ratings of the written record. The study did collect one patient-reported outcome — whether patients said they were still not feeling better at follow-up — at 3.8% with the AI against 4.3% without, a difference that did not reach significance. And the diagnostic and treatment effects were null until Penda layered on peer champions, one-on-one coaching and leaderboards; only the history and investigation effects survived the induction period. The measured AI benefit depended on exactly the kind of human support layer the 2025 funding collapse removed elsewhere.
Then, eleven months later, the harder test. A registered, Gates-funded pragmatic cluster-randomized trial of the same tool across the same Penda network — 16 clinics, 103 clinical officers, 9,691 patients — was published in Nature Medicine. On the patient outcome it was null: 14-day treatment failure 2.2% with AI against 2.0% without (aOR 0.77, 95% CI 0.55–1.08). On documentation it was positive: appropriate diagnosis and comprehensiveness both improved. That is not evidence that AI does nothing: a confidence interval that wide leaves the question open. It is evidence that the shape of the EMR result repeated itself in a randomized design, under the best conditions anyone has arranged for it, with the vendor’s own product and a well-run clinic chain.
A third paper from the same deployment rarely makes it into the summaries. A retrospective physician-panel review of 1,469 encounters found actively harmful AI recommendations in 7.8% of them, 37 of those major — and that clinicians fully adopted 22% of those harmful recommendations and partially adopted a further 37%, leaving 42% unacted on. The same review found the effect running the other way too: in 8.0% of encounters the assistant caught and mitigated a risk in the clinician’s own notes, and where its advice was followed at all, beneficial recommendations were taken up roughly twice as often as harmful ones (118 against 67). What the authors flag is the asymmetry in the other direction — 362 separate instances where beneficial guidance was offered and ignored. Its own authors put mitigations on the record: no adverse patient outcomes were found on follow-up, ten of the 37 major flags came from a single evaluator, and a post-study audit judged many to be judgement calls. It is also not independent — same lead authors as the trial, Penda co-authors, two holding Penda stock options. Treat it as a signal worth watching, and one that belongs beside the headline result.
A fourth paper from the same clinics answers the question the other three skip: whether clinicians use the thing at all. Across 258,106 clinical episodes, they opened the assistant in 21.7% — adoption climbing from 4% to 47% over the study — and used it least when they felt confident, which the authors say leaves clinician-initiated querying “liable to overconfidence-related underutilisation”. That is the deployed-is-not-used pattern from the EMR era reappearing intact, one software generation later.
The scribe question is at an earlier stage. A preprint from Groote Schuur Hospital’s neurosurgery division in Cape Town scored ambient-AI notes against contemporaneous handwritten ones across 49 encounters and found the AI notes far better — mean SOAP quality 4.9 against 2.9, with moderate-to-severe error rates at least fivefold higher in the handwritten notes and clinically significant impact in 38.8% of handwritten notes against 2.0% of AI ones. Raters agreed strongly with each other. They were also not blinded, the scribe vendor has a co-author, the encounter audio serves as both the AI’s only input and the scoring reference, and it is 49 encounters in one division that has not yet been peer-reviewed. Handwriting is a low bar, and clearing it is worth something in a system drowning in paper. But it says nothing about patients, and neither does anything else: as of August 2026 we could find no AI scribe study anywhere in the world — not in Africa, not in the far larger US and European deployments — tested against a patient health outcome.
| Finding | Effect | Design — read this column | Setting | Source |
|---|---|---|---|---|
| AI clinical decision support improved the record, not the patient outcome | 14-day treatment failure 2.2% vs 2.0% (aOR 0.77, 95% CI 0.55–1.08, P=0.13) — no evidence of benefit, not evidence of none. Documentation secondaries positive: appropriate diagnosis aOR 1.74, comprehensive documentation aOR 1.68, appropriate treatment plan aOR 1.71 (all P<0.001, from 2,000 reviewed encounters). No safety signal. | Pragmatic cluster RCT, 16 clinics, 103 clinical officers, 9,691 patients, Apr–Jul 2025. Registered (PACTR202502499779176), Gates-funded, PATH-sponsored, no OpenAI role. Clinicians randomized within shared facilities — any contamination biases toward the null. | Kenya, Penda Health primary care | source |
| LLM safety-net lowered physician-rated errors in the written note | −16% diagnostic (95% CI 6.9–24.2) and −13% treatment (6.8–18.3) relative error reduction. Both were null before an intensive coaching layer was added (treatment 4.3%, −3.0 to 11.1); history and investigations effects survived without it. Patient-reported 'not feeling better' at 8 days was null (3.8% vs 4.3%, ~39% response). | Quality-improvement study the paper calls pragmatic cluster-assigned: clinicians randomly assigned to tool access within each of 15 clinics (57 with / 49 without), unregistered. Errors are physician Likert ratings of a ~14% random sample (5,666 visits rated), headline figures main-period only — not 'across 39,849 visits', a total that does not reconcile with the paper's own flow diagram (39,579). Inter-rater agreement fair (Fleiss κ 0.22–0.29); 4,279 visits rated by a single physician. AI-arm notes were longer. OpenAI funded the study and was involved in analysis and reporting. | Kenya, Penda Health | source |
| The same deployment produced harmful recommendations, and most of them were acted on | Actively harmful recommendations in 115 of 1,469 encounters (7.8%, 95% CI 6.5–9.3; 37 major, 78 minor), 67 reaching the final documentation. Of the 115, clinicians fully adopted 25 (22%) and partially adopted 42 (37%); 48 (42%) were not acted on. Where advice was followed at all, beneficial recommendations were taken up about twice as often as harmful ones (118 vs 67), and 362 beneficial recommendations were offered but ignored. Hallucinations 3.4%. Clinician-side risk fully mitigated in 8.0%. Follow-up confirmed no resulting adverse patient outcomes; 10 of the 37 major flags came from one evaluator, and a post-study audit judged many to be judgement calls. | Retrospective physician-panel review of 1,469 records, 16 clinics, Jul–Sep 2024. Not independent: same lead authors as the randomized trial above, Penda co-authors, two holding Penda stock options. | Kenya, Penda Health | source |
| Clinicians used the AI assistant in about a fifth of consultations, and least when they felt confident | 56,050 of 258,106 clinical episodes (21.7%) over Feb–Oct 2024; adoption rose 4%→47% over the 8-month period; clinicians left feedback on 31% of episodes and 99.5% of it was positive. The authors conclude clinician-initiated querying 'is liable to overconfidence-related underutilisation.' | Mixed methods: CDSS metadata from all consultations plus 42 staff via interviews and focus groups; PATH-led with Penda co-authors | Kenya, Penda Health, 16 facilities | source |
One further mortality number belongs on the record precisely because it cuts against the technology. A Nigerian trial of an AI digital stethoscope in obstetric care — SPEC-AI, 1,195 women across six teaching hospitals — reported an unexplained excess of all-cause deaths in the AI arm, 12 against 3, hazard ratio 4.20 (95% CI 1.18–14.87). It is a secondary, unpowered finding with no adjustment for multiple comparisons, cardiovascular mortality was null, and the trial’s own primary endpoint was how much disease was found rather than how much occurred, which is why it is absent from the tables above. It should not be read as evidence of harm — and it should not be quietly dropped either.
How to read the next claim
The pattern across every strand of this evidence is consistent enough to use as a test. Ask what moved. In this literature, process and documentation measures are the easiest things to move — and the easiest to measure; clinical-action outcomes move inconsistently; and hard patient outcomes are not as rare as the field implies — roughly a dozen African trials have measured death, though only two were powered to detect a difference in it. What is rare is a positive one. No trial powered to detect a mortality difference has found one, and six point estimates favour the control arm. Two results point the other way and both deserve their weight: a Zanzibar messaging trial found fewer perinatal deaths as a secondary outcome, and a Kenyan TB adherence platform moved a composite containing death — where the largest effect was in the arm with a trained human support team attached. The one study that found fewer deaths outright, in Malawi, was not randomized. Even the documentation gains are not uniform, as the Rwandan and Kenyan assessments above show: at national scale the electronic record was sometimes worse than the paper it replaced. A systematic review of 25 randomized eHealth trials across ten sub-Saharan African countries found the same split — behaviour changed (OR 1.21, 95% CI 1.05–1.40, pooled over 11 studies) while biological outcomes stayed undetected (OR 1.17, 0.89–1.54, pooled over 6). When a vendor or a funder cites a documentation improvement as though it were a health improvement, that is the gap they are stepping over.
Ask who does the work the tool depends on. Malawi’s mortality result required tracers who went and found people. Kenya’s loss-to-follow-up result required community social workers making calls and home visits — and even then only 23–31% of lapsed patients were relinked. Penda’s AI gains required coaches and champions. Every result in this file that touched a patient had a person attached to it. The tracing and data-entry workforce behind the Malawi and Kenyan results was donor-funded, and that funding is what collapsed in 2025; Penda’s coaching layer was paid for by a private clinic chain, which is its own kind of fragility.
Ask whether the outcome is the thing or the record of the thing. The Kenyan alert trial measured documented loss to follow-up while the alerts explicitly instructed staff to revise that documentation; an unknown share of the improvement is reclassification. When routine records misclassify a third of outcomes, any study whose endpoint is drawn from those records inherits the error.
Ask what else the same paper reports. This is the cheapest check available and it worked every time here: the Malawi paper contains a null on care at the visit, the Kenyan trial’s parent publication contains a marginal clinical result its daughter paper omits, the Malawi programme report contains an admission that the reports came out wrong, and the Penda deployment contains a harms audit. In each case the inconvenient finding sat inside the source already being cited.
Five questions do most of the work. They are what this review applied to every study in the tables above, and they are as useful on a press release as on a paper.
- What actually moved — a process measure, a clinical action, or a patient outcome? Documentation and process metrics are the easiest things to move and to measure. Name which of the three the headline number belongs to, and check whether a patient outcome was measured at all.
- Who does the work the tool depends on? Look for the humans in the mechanism — tracers, data clerks, coaches, community workers. If the effect required a paid support layer, the result does not transfer to a setting without one.
- Is the outcome the thing, or the record of the thing? An intervention that instructs staff to update a record, measured by that record, may be capturing reclassification rather than real change. Ask how the endpoint was ascertained and how often routine records are wrong.
- What else does the same source report? The cheapest check available. Look inside the cited paper for a null, a harm, a failure or a caveat that the headline skipped — it is usually there.
- Is there a later or better-designed study of the same product? An early uncontrolled or unregistered result is often superseded within months. Check for a registered trial, an independent evaluation, or a replication before treating the first number as settled.
None of this says the systems were a waste. An 11–20% reduction in annual deaths, reaching 28% by year five, is a large effect, and clinics that can find their lapsed patients save lives. It says something narrower and more useful: these systems earn their keep as instruments of follow-up and accountability, staffed by people. Improving patient outcomes through better bedside decisions is the thing they have not yet been shown to do. The AI wave is currently producing evidence of exactly the same kind, which is a reason to fund the tracing and the coaching, and a reason to be sceptical of anyone who tells you the model itself is the intervention.
The five questions, run on one claim
Take the headline from that first Penda study. It comes from a real study, the number is accurately reported, and the tool is a serious one.
Claim under test “AI Consult reduced diagnostic errors by 16% and treatment errors by 13% at Penda Health, Kenya.”
- What moved? A process measure. The “errors” are physician ratings of the written clinical record.
- Who does the work? Penda layered peer champions, one-on-one coaching and leaderboards on top of the tool. The diagnostic and treatment effects were null until that human support arrived — so the number describes a tool plus a coaching programme around it.
- The thing, or the record of it? The record. Raters scored notes, and the AI writes into the notes — so a tool that improves documentation will move the measured outcome whether or not care changed.
- What else does the source say? The same paper reports a patient-reported outcome — whether patients were still not feeling better at follow-up — at 3.8% with the AI and 4.3% without, a difference that did not reach significance. A companion audit of the deployment found actively harmful recommendations in 7.8% of encounters, and clinicians adopting harmful advice more often than beneficial advice. A third paper found clinicians opened the assistant in only 21.7% of episodes.
- Is there a better study? Yes — and it supersedes the first. A registered, independently funded cluster-randomized trial of the same tool, at the same clinics eleven months later, found no reduction in 14-day treatment failure (2.2% vs 2.0%) while confirming the documentation gains.
What survives A real, replicated improvement in the medical record, dependent on a human coaching layer — and no demonstrated patient benefit. Note the honest limit in the other direction: the trial’s confidence interval is wide enough that a real effect could exist and simply have gone undetected.
Evaluate a claim yourself
Take the five questions to your own AI
Paste a study, evaluation, press release, or vendor claim below. It turns the five questions into a prompt that makes any assistant show its working — name what moved, find the humans in the mechanism, and keep what was measured apart from what was asserted. Copy it into ChatGPT, Claude, or Gemini.
You are evaluating an impact claim about {{SUBJECT}}, using the "Ground Truth" method (groundtruth.health). I will give you a study, evaluation, press release, or vendor claim. {{MODE}}
1. What actually moved — a process measure, a clinical action, or a patient outcome? Documentation and process metrics are the easiest things to move and to measure. Name which of the three the headline number belongs to, and check whether a patient outcome was measured at all.
2. Who does the work the tool depends on? Look for the humans in the mechanism — tracers, data clerks, coaches, community workers. If the effect required a paid support layer, the result does not transfer to a setting without one.
3. Is the outcome the thing, or the record of the thing? An intervention that instructs staff to update a record, measured by that record, may be capturing reclassification rather than real change. Ask how the endpoint was ascertained and how often routine records are wrong.
4. What else does the same source report? The cheapest check available. Look inside the cited paper for a null, a harm, a failure or a caveat that the headline skipped — it is usually there.
5. Is there a later or better-designed study of the same product? An early uncontrolled or unregistered result is often superseded within months. Check for a registered trial, an independent evaluation, or a replication before treating the first number as settled.
Prefer primary sources; if a claim can only be traced to a press release or blog post, treat the number as unverified. Distinguish "no evidence of benefit" (a study that could not detect an effect) from "evidence of no benefit"; a wide confidence interval means the first, and leaves the question open. Do not fill gaps with assumptions — say "not stated" wherever the source is silent.
CLAIM / STUDY TO EVALUATE:
{{CLAIM}}
Ground Truth doesn’t run the model for you — deliberately. The method is ours; the judgement stays yours.
Sources and method
Every study here was pulled to its primary source and read for its design as well as its headline number, and each effect size was checked against that source independently. Where a study cannot bear the weight commonly placed on it, the table says so in the design column. Facts current as of August 14, 2026. Corrections are published in the open corrections log.
Disclosures & provenance
- Published
- 27 Aug 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Corrections
- None to date. Corrections log → · Challenge this analysis