In October 2025, Butterfly Network announced that clinicians in Malawi and Uganda were using a new kind of ultrasound AI. A health worker would move a handheld probe across a pregnant patient’s abdomen in a fixed series of “blind sweeps”: no hunt for the perfect image, no manual measurement of the fetal head, abdomen or femur. The software would turn the resulting video into an estimate of gestational age.

The company described the tool as requiring “no image interpretation or provider training”. It said the estimate could help clinicians choose the right time for life-saving medicines and procedures. When the US Food and Drug Administration cleared the product in March 2026, Butterfly said it could guide timely care and help reduce maternal and fetal mortality.

Underneath that story is a narrower scientific claim: a novice using AI can estimate how far along a pregnancy is about as accurately as a credentialed sonographer performing fetal biometry. That claim matters. Gestational age determines when clinicians screen, treat and deliver; in settings without reliable menstrual dates or access to sonographers, a useful estimate can change what care is possible.

It is also the kind of comparison that often falls apart when the two sides of the test are reconstructed. This one does not. The central result survives.

The boundary around it is where the story changes.

The number survives

The strongest evidence comes from a prospective 2024 diagnostic-accuracy study in JAMA. Researchers enrolled 400 pregnant participants at two sites in Lusaka, Zambia, and one in Chapel Hill, North Carolina. Before 14 weeks, a credentialed sonographer established each pregnancy’s reference age using transvaginal crown-rump length — the most accurate clinical dating method available early in pregnancy.

Each participant returned at a randomly assigned week. At the same visit, a novice collected blind sweeps using an AI-enabled Butterfly iQ+ and a credentialed sonographer measured the fetal head, abdomen and femur on a high-specification GE machine. Both were blinded to the early dating result and to each other’s estimate. The AI result was hidden from the novice, preventing it from influencing the expert examination or the patient’s care.

The primary window was 14 weeks through 27 weeks and 6 days. Among the 399 participants for whom both methods returned a result, the AI workflow’s mean absolute error — the average distance between its estimate and the early-pregnancy reference — was 3.19 days. Expert biometry’s was 3.03 days. The difference was 0.16 day, with a 95% confidence interval from −0.14 to +0.45 day.

The prespecified statistical plan defined the methods as equivalent if that entire interval stayed between −2 and +2 days. It did so easily. The paper says the two-day boundary was set through consultation with experts in North Carolina and Zambia; it does not tie that threshold to a measured clinical consequence. As a descriptive, post hoc observation — not a second prespecified equivalence test — the reported interval would also fit inside ±0.5 day.

Equivalence here applies to population-average absolute error. It does not mean that the AI and sonographer gave the same answer for every patient. The distribution nevertheless supports the average. The AI estimate was under seven days from the reference in 90.7% of cases, compared with 92.5% for expert biometry. Both methods were under 14 days in 99.8% of cases — one miss for each method — and root-mean-square error was also similar.

The primary finding was supported by more than mean absolute error alone: the threshold and root-mean-square-error analyses pointed in the same direction. On the question asked and in the enrolled population, the AI was about as accurate as expert biometry.

The 2024 article reports no conflicts of interest. The Gates Foundation funded the work; Butterfly donated probes and worked with investigators to integrate an encrypted version of the model. The article says neither organization had a role in study design or conduct, data collection or analysis, manuscript preparation, or the publication decision. Its Global Access disclosure also identifies international patent publication WO/2023/122326.

Dating error widens with gestation, for both methods
Mean absolute gestational-age error for blind-sweep AI and expert fetal biometry in three gestational windows A paired dot chart. At 14 to 27 weeks, AI error was 3.19 days and expert biometry error was 3.03 days. At 28 to 36 weeks, AI error was 6.07 days and expert error was 7.12 days. At 37 to 40 weeks, AI error was 11.54 days and expert error was 9.10 days.
Source: Stringer et al., JAMA (2024), Table 2 and Supplement 2.
Data table
Evaluation windowBlind-sweep AI, MAEExpert biometry, MAEPaired differenceStudy interpretation
14–27 weeks, primary3.19 days3.03 days+0.16 days (95% CI, −0.14 to +0.45)Equivalent
28–36 weeks, secondary6.07 days7.12 days−1.06 days (95% CI, −1.72 to −0.40)Equivalent; prespecified secondary result favored AI
37–40 weeks, tertiary11.54 days9.10 days+2.43 days (95% CI, +1.19 to +3.68)Equivalence not established; AI had higher MAE
Read before citing
  • Equivalence is a population-average result, not agreement for every pregnancy.
  • The primary claim covers 14 through 27 weeks and 6 days; the 28–36-week result was prespecified but secondary.
  • At 37–40 weeks, equivalence was not established and the AI systematically underdated pregnancies. The authors advised against term use.
  • Later ultrasound dating is less precise for both methods because fetal size increasingly reflects growth as well as age.

“No prior sonography training” is not “no training”

The word novice carries much of the claim, and it is accurate with one important translation.

The study used 12 novice operators. Seven in Zambia were research nurses or midwives. Five in North Carolina were research assistants with bachelor’s degrees and no clinical role. The paper described all 12 as having no prior sonography training. Before collecting study scans, however, they received a full day of instruction covering the software, patient positioning, gel, probe orientation, pressure and sweep collection. Half of that day was hands-on practice with patients under the supervision of an experienced sonographer.

The app did remove a large part of the traditional skill burden. Operators did not interpret a live ultrasound image. An animation directed a sequence of 10-second sweeps across the abdomen. The operator first measured fundal height, which the workflow used to choose the number of sweeps and configure probe depth and gain. The model processed ultrasound frames; the paper does not describe the fundal-height value itself as a direct predictive input, although it shaped sweep count, depth and gain.

The workflow also contained a quality fail-safe. If the model found no frame with a sufficiently strong attention score, it asked the user to repeat the acquisition. Only after five failed collections would it return no result. Final availability was excellent: 399 estimates from 400 primary-window examinations. But the paper did not report first-pass success, the average number of attempts or the time added by retries; it did state that novice users could not consult study sonographers while using the tool.

The paper also did not report results by operator or an operator-clustered analysis, so consistency across all 12 users remains uncertain. That is an operator-generalizability gap, not evidence that the pooled result is false.

The defensible reading is therefore stronger than “only sonographers can do this” and narrower than “no training required.” The evidence supports use by people with no prior sonography training after brief, task-specific instruction. It does not test someone opening the app and scanning successfully with no preparation.

That distinction matters because Butterfly’s Africa launch material uses the broader wording. A company could reasonably say that “no provider training” meant no formal sonography education, rather than no product onboarding. But those are not the same promise. The studies establish the first. We found no published study that evaluated the second.

Selected pregnancies, one output

The participants were not a cross-section of every patient for whom late pregnancy dating might be useful. They had viable singleton pregnancies, BMI below 40 and no known major fetal anomaly. They had presented early enough to receive a first-trimester dating scan and intended to remain nearby for follow-up. The Zambia sites were urban research clinics. People with multiples, known anomalies or conditions likely to make participation unsafe or complicate interpretation were excluded.

Needing an early reference is not a flaw in the accuracy test. Without a reliable early date, there is no good way to know how wrong a later estimate is, and the study teams kept that date hidden from both novice and expert operators. The trade-off is generalizability: performance remains unestablished for people with BMI of 40 or above, multiples, known major anomalies and other excluded conditions; the cohort’s early presentation and ability to return may also differ from the intended late-entry population.

One prespecified subgroup merits attention. Among 97 participants with BMI of at least 30, AI mean absolute error was 3.78 days versus 3.07 for biometry, a difference of 0.70 day (95% CI, 0.07 to 1.33). That interval still fit inside the ±2-day equivalence margin, but subgroup comparisons were not adjusted for multiplicity, and no participant had BMI of 40 or higher.

The 2022–23 validation cohort was entirely separate from the 2018–21 development data, a real strength. But the same research group tested the model in the same two countries, making this prospective temporal validation rather than independent-developer or broad geographic replication.

Later pregnancy reveals the tool’s other boundary. In a prespecified secondary analysis at 28–36 weeks, the AI’s average error was 6.07 days, versus 7.12 days for the sonographer. The AI met the equivalence criterion and had lower average error, a prespecified secondary result. At 37–40 weeks, among 175 participants still pregnant and assessed at that visit, both methods had larger errors and the AI had higher mean absolute error: 11.54 days versus 9.10. Its mean signed error had the same magnitude as its mean absolute error to two decimals, indicating that almost all error pointed in one direction — toward underdating. Only 28.0% of AI estimates were under seven days from the reference, compared with 46.3% for biometry. Equivalence was not established, and the authors advised against term use.

That pattern is not uniquely an AI problem. Clinical guidance from ACOG calls third-trimester ultrasound the least reliable way to date a pregnancy because fetal size increasingly reflects constitutional and pathological growth, not age alone. The correct conclusion is not that the AI “collapses” while ordinary ultrasound stays precise. The 2024 equivalence result extends through 36 weeks and 6 days; its 37–40-week analysis did not establish equivalence. FDA later cleared the commercial product for pregnancies presumed to be 16–37 weeks on a separate submission, so the research window and cleared range should not be treated as identical.

It also stops at dating. The 2024 Butterfly GA study did not test for ectopic pregnancy, multiples, fetal anomalies, placental location, fetal growth, amniotic fluid or malpresentation. The World Health Organization’s recommendation for one imaging ultrasound before 24 weeks includes gestational age, but it is broader than gestational age alone. A blind-sweep calculator may supply an important component of that recommended scan. It is not equivalent to a complete obstetric examination.

Two model lineages, shared roots

The story became easier to misread in July 2026, when a second team published another strong blind-sweep result.

Both research implementations draw on ultrasound data collected through the Fetal Age Machine Learning Initiative, or FAMLI, in Chapel Hill and Lusaka. But the published record describes at least two research implementations, not repeated prospective tests of one documented frozen model.

Pokaprakarn and colleagues described the model that Stringer and colleagues later evaluated prospectively on Butterfly iQ+ hardware. Gomes and colleagues described the Google model that Willis and colleagues adapted and evaluated in 2026. The papers describe different architectures and implementations; they do not establish that the two studies used the same weights or software.

FDA and Butterfly identify the UNC/FAMLI work as the research basis for Butterfly’s commercial tool. Public sources do not establish that the FDA-cleared commercial model is technically identical to the frozen research model evaluated in 2024.

That distinction matters because the new 2026 JAMA Network Open study evaluates the Google lineage, not the commercial Butterfly tool.

One research program, two published model lineages
Published-evidence map for two FAMLI blind-sweep ultrasound model lineages and the separate commercial Butterfly tool The FAMLI data program in Chapel Hill and Lusaka supports two published branches. One runs from the Pokaprakarn 2022 research model to the Stringer 2024 prospective evaluation on Butterfly iQ Plus. The other runs from the Gomes 2022 Google model through Clarius C3 adaptation to the Willis 2026 Chicago and Nairobi study. A separate dashed box identifies the FDA-cleared Butterfly commercial tool; no arrow claims direct frozen-model continuity from the 2024 study. PUBLISHED-EVIDENCE MAP · NOT CODE GENEALOGY FAMLI ultrasound data program Chapel Hill, North Carolina + Lusaka, Zambia Pokaprakarn et al. 2022 UNC/FAMLI research model published development evidence Gomes et al. 2022 Google research model different published architecture Stringer et al. 2024 Prospective evaluation on Butterfly iQ+ Zambia + North Carolina Clarius C3 adaptation 180 Chicago exams · 300,000 steps Willis et al. 2026 Chicago + site-held-out Nairobi Butterfly Gestational Age Tool · FDA-cleared 2026 UNC/FAMLI research basis; exact frozen-model continuity is not public
Evidence map as a table
Published branch or productDevelopment evidenceProspective evaluation or statusWhat the map does not establish
Pokaprakarn/Stringer research branchPokaprakarn et al. 2022, FAMLI dataStringer et al. 2024 on Butterfly iQ+ in Zambia and North CarolinaTechnical identity with the current commercial model
Gomes/Willis Google research branchGomes et al. 2022, FAMLI dataClarius C3 adaptation, then Willis et al. 2026 in Chicago and NairobiReplication of the Butterfly product
Butterfly Gestational Age ToolFDA and Butterfly cite UNC/FAMLI as the research basisFDA-cleared in 2026 under K252148Direct continuity from the frozen 2024 research implementation
How to read this map
  • It maps published evidence, not ownership, patent scope or source-code genealogy.
  • The two research branches share FAMLI data-program roots and overlapping collaborators, but the papers describe different architectures and do not document shared frozen weights.
  • The 2026 Google study is class-level corroboration for blind-sweep dating. It is not replication of the commercial Butterfly tool.
  • The commercial product is intentionally shown without a direct arrow from the 2024 study because exact frozen-model continuity is not public.

The Google study reached the same performance range. In the 385-participant primary set, examined from 16 weeks through 36 weeks and 6 days, the model’s mean absolute error was 4.2 days (95% CI, 3.8–4.6), versus 4.5 days (95% CI, 4.1–4.9) for Hadlock biometry. The authors reported noninferiority under a one-day margin (P < .001). Site-level point estimates were 4.1 versus 4.6 days in Chicago and 4.3 versus 4.3 days in Nairobi.

The point estimates corroborate the blind-sweep approach, but the generalization claim has layers. Researchers adapted the model to the Clarius C3 probe using 180 examinations from 120 Chicago participants, then trained for 300,000 optimization steps with batches divided evenly between the Clarius and original FAMLI data. Both evaluation sites used the same Clarius C3 hardware. Chicago’s primary participants were held out from adaptation, but they came from the adaptation institution; Nairobi was the only site-held-out cohort and contributed no fine-tuning data. The study therefore supports transfer to a new site and new operators after probe-specific adaptation — not zero-shot transfer across devices or arbitrary settings.

Reference dating was not uniform across sites. Every Chicago participant had a dating scan before 14 weeks using crown-rump length. In Nairobi, 119 participants had a scan before 14 weeks, while 74 — 38.3% — were dated between 14 and 22 weeks using fetal biometry, a less precise reference. The reported model error was similar in those Nairobi subgroups, 4.2 versus 4.4 days (P = .98), which is reassuring but does not make the reference standards equivalent. The authors also report that Nairobi’s standard-care sonographers were not blinded to the known dating age, a design feature they cite when interpreting the late-term comparator.

The study used five novice operators: two in Chicago, who received a handout, verbal instruction and a demonstration, and three in Nairobi, who received about six hours of practical training. That supports use across more than one operator and setting, but five people are a thin basis for claims of operator-independent performance.

The largest unresolved question is who reached the primary analysis. Of 1,050 participants deemed gestational-age eligible, 120 were used for adaptation, 63 were beyond the primary window, and 482 were excluded for “no valid index test.” After setting aside adaptation and late-term records, that undefined exclusion affected 437 of 629 otherwise eligible Chicago participants (69.5%) and 45 of 238 in Nairobi (18.9%).

The paper does not define the category well enough to distinguish missing comparator scans, acquisition or processing problems, predeployment records, or other causes. Those records cannot responsibly be relabeled as AI failures. But the unexplained and strongly site-dependent attrition prevents readers from estimating case-level coverage or judging how representative the analyzed set is. The 4.2-day result applies to the 385-participant primary set, not everyone entering the study workflow.

The paper does report acquisition-level information for the analyzed set: 4.9% of individual sweeps were rejected, and examinations contained an average of 7.2 attempted sweeps. Those numbers do not resolve how often an eligible patient failed to reach a final case-level estimate.

The rounded point estimates are compatible with noninferiority, but the formal analysis is not fully reproducible from the report. Each participant contributed AI and Hadlock errors, making the comparison paired; separate confidence intervals for each method do not supply the confidence interval for the paired difference. The paper does not report that difference at full precision, its confidence interval or a primary test statistic. Its methods list a t test or Welch t test and, for nonnormal data, a Wilcoxon or Mann–Whitney test, without stating which produced the primary P value; some of those tests ordinarily assume independent rather than paired samples. That ambiguity is an author query, not proof that the analysis was wrong.

The cohort size was determined by available enrollment rather than a reported noninferiority power calculation. The article and its supplements provide no registration identifier, protocol or statistical-analysis plan, participant-level abstention rate, or primary tail metric such as the percentage missing by more than seven or 14 days.

Google supported the study through grants to Jacaranda Health and Northwestern University, and the Gates Foundation partially funded Northwestern. Google-employed authors participated in study design and conduct, management, analysis, interpretation, manuscript preparation and the publication decision; Google did not collect the data. Nairobi supplies geographically and institutionally held-out evaluation data, but this was not replication by an unaffiliated developer or investigator group.

The gestational window is also a hard boundary for this model. At 37–40 weeks, the paired Chicago sample was small (n = 9), but the model’s mean absolute error rose to 9.6 days. In Nairobi (n = 54), it rose to 12.0 days, versus 5.0 days for Hadlock. The model’s signed errors were positive by approximately the same amounts, indicating systematic underdating rather than symmetric scatter. A larger model-only Chicago set (n = 114) showed the same direction, with 7.8 days of absolute error and 7.7 days of underdating. Nairobi’s unblinded comparator complicates the head-to-head comparison, but not the model’s directional error against the dating reference. This does not invalidate the 16-to-36-week result; it prevents extending that result to term.

None of those issues disproves the narrow result. They keep it from carrying more weight than it earned. The 2026 study corroborates a class of approach within a bounded window. It does not independently validate Butterfly’s product.

What FDA cleared — and what it did not

In March 2026, FDA cleared the Butterfly Gestational Age Tool through the 510(k) pathway. The indication for use is precise: an estimate for a singleton intrauterine pregnancy presumed to be between 16 and 37 weeks, using an iQ+ or iQ3 probe, for qualified and trained healthcare professionals. The output is adjunctive and is “not intended to be used for prenatal management and/or delivery planning”.

That is a meaningful regulatory clearance, and it matches much of the research boundary. It does not establish a mortality benefit, and the cleared indication is for qualified and trained healthcare professionals — not zero-instruction use.

The manufacturer’s public 510(k) summary reports a US evaluation of 111 participants scanned by 13 trained healthcare professionals — seven physicians and six sonographers. It compared the tool and conventional biometry against reported last menstrual period and compared the two ultrasound estimates directly. On the submitted record, FDA determined the device substantially equivalent for the stated indication.

That evaluation used a weaker dating reference than the 2024 study’s first-trimester crown-rump length. Two of the four study locations were Butterfly offices and accounted for 94 of the 111 participants; the public summary does not report the operators’ affiliations. It did not test nurses, midwives or nonclinical research assistants like the novice operators in the published studies. The label does not claim a replacement for standard biometry.

Butterfly’s current user manual makes the boundary plainer. It says the tool is not a replacement for ultrasound-based biometry, should not change a gestational age or due date established earlier, and may be inaccurate with multiple fetuses, BMI above 40, known major anomalies or maternal conditions that affect fetal size or growth.

The company’s public story travels further. Its US clearance announcement says the estimate can guide timely care, enable faster emergency decisions and help reduce mortality. Its Africa launch says no provider training is required and that an accurate estimate helps clinicians select life-saving medicines and procedures.

Those are plausible pathways from better dating to better care. They are not endpoints tested by the diagnostic studies, whose AI estimates were deliberately hidden from care teams. And the US label does not govern product use in Malawi or Uganda, so the comparison should not be mistaken for an allegation of unlawful marketing. The narrower conclusion is enough: the promotional language describes a broader hoped-for care pathway than the published studies or US indication establish.

The missing rungs between accuracy and impact

Butterfly reports that the GA tool is deployed in Malawi and Uganda. As of July 14, 2026, we found no public GA-tool-specific implementation report stating how many facilities, clinicians or scans were involved, how often the system produced a usable result, its uptime, what referrals or treatments changed, or whether outcomes differed.

That does not mean the deployment failed. It means the deployment is not publicly evaluable.

The evidence ladder
Evidence rungWhat is public?Status
Can the algorithm estimate gestational age accurately in selected pregnancies?Two model lineages, prospective cohorts, expert comparisonSupported in the analyzed cohorts before term
Can novices collect usable sweeps after brief instruction?2024 and 2026 research workflowsSupported in the 2024 study; 2026 case-level coverage is not reported
Does the tool work reliably at routine scale?No public GA-tool-specific facility, coverage, retry, uptime or failure report identified as of July 14, 2026Not publicly established
Does showing the estimate change clinical decisions appropriately?AI output hidden in the principal studiesNot tested
Does deployment improve maternal or newborn outcomes?No controlled GA-tool impact evaluation identified as of July 14, 2026Not tested

Cost has the same missing middle. The 2024 paper says UNC will make the machine-learning model available at no cost in low- and middle-income countries under Gates Foundation Global Access requirements. The model is not the complete system: a compatible probe and mobile device, gel, charging, maintenance, support and training still have costs. As of July 14, 2026, Butterfly listed individual US prices of $2,699 for iQ+ and $3,899 for iQ3. Butterfly says the GA Tool requires an active Advanced membership; its individual US pricing page listed Advanced at $420 per year or $1,500 with software access guaranteed for five years. These figures illustrate why no-cost model access does not disclose total delivered-system cost; they are not Malawi or Uganda procurement prices.

Butterfly cites preliminary company-and-partner analyses from its separate 1,000 Probe Partnership in Kenya and South Africa; the company’s announcement itself cautions that other interventions may have contributed to the observed changes. That broader obstetric point-of-care ultrasound program involved substantial training and did not isolate this gestational-age algorithm, so its outcomes cannot establish the GA Tool’s effect.

The distinction is not academic. The studies establish measurement accuracy. To establish impact, the result has to be shown to a clinician, interpreted in the context of a real patient, connected to a functioning referral and treatment system, and used without creating new harms through false reassurance or misdating. Each step may work. None follows automatically from a four-day mean error.

What would move this rating

The next decisive study is not another comparison of average dating error. It is a prospective implementation study that reports who can use the tool, after what instruction, at what total cost, how often it produces a usable result, what clinicians do differently because of it, and whether patients fare better.

The bottom line

The central technical result holds up: in selected singleton pregnancies before term, novices described as having no prior sonography training — after brief, tool-specific instruction — used blind-sweep AI to estimate gestational age about as accurately as credentialed sonographers performing fetal biometry. The best study was prospective, paired, blinded and preregistered. A second model lineage produced a similar error estimate in its analyzed Chicago and Nairobi cohort, but unresolved exclusions and thinner statistical reporting make it corroboration of the approach, not independent validation of Butterfly’s product. Butterfly’s version is FDA-cleared for a narrow adjunctive indication. That is a useful achievement. It does not establish no-instruction use, equivalence to a complete obstetric scan, reliable operation at scale, better clinical decisions, or lower maternal and fetal mortality. The research demonstrated an access-enabling measurement system; it did not demonstrate expanded access or improved outcomes.

Primary studies and regulatory sources: Stringer et al., JAMA (2024) · NCT05433519 statistical analysis plan · Pokaprakarn et al., NEJM Evidence (2022) · Gomes et al., Communications Medicine (2022) · Willis et al., JAMA Network Open (2026) · FDA 510(k) record K252148 · FDA clearance letter, indication and summary.
Claim as circulated: Butterfly Africa launch · Butterfly FDA-clearance announcement.
Product and practice context: Butterfly user manual · Butterfly US pricing · Butterfly terms of use · WHO antenatal-ultrasound recommendation · ACOG, “Methods for Estimating the Due Date”.
Ground Truth is independent of, and unaffiliated with, Butterfly Network, Google, UNC, Clarius and the study funders. Corrections are made in public — see our policy.

Disclosures & provenance

Published
14 Jul 2026
Author
The Ground Truth editor. Editorial standard →
Funding
Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
Rating
Holds up with conditions (4/5), rubric v1.0, rated 14 Jul 2026. Full rating card →
Corrections
None to date. Corrections log → · Challenge this analysis