One person can appear in two study counts without becoming two people.

On 25 August, AEYE Health and academic collaborators published the results of three pivotal studies of an autonomous diabetic-retinopathy screening system. The paper’s title says the studies included “over 1,200 patients.” Its three displayed cohort counts add to 1,213: 531 in AEYE-1, 317 in AEYE-2 and 365 in AEYE-3. The 1,210 total can be reconstructed only by mixing cohort stages: 531 screened in AEYE-1, 317 enrolled in AEYE-2 and 362 enrolled in AEYE-3. The arithmetic is recoverable, but the denominator is not stage-consistent.

But the methods say all 317 people in AEYE-2 were a subset of the 531 in AEYE-1, tested again with a different camera. Add the two independent cohorts—not the repeated camera substudy—and the maximum number of unique people screened is 896. Using the displayed cohort sum as the denominator, at least 317 of the 1,213 camera-study entries, 26.1%, represent participants counted a second time.

The diagnostic result survives. The denominator needs a cohort-lineage qualifier.

The paper was funded by AEYE Health. Its disclosures list six authors as company employees and one additional author as a board member; the paper says the funder was not involved in study design, data collection, analysis, interpretation or the decision to publish. That does not negate the results, but it raises the importance of a reproducible denominator and clear cohort accounting.

We reconstructed all three diagnostic matrices and reproduced the reported sensitivity and specificity. The cohort chart below shows the exact counts and reconstructed imageability fractions. Those fractions are Ground Truth reconstructions—not printed cohort-flow counts or all-screened technical-yield fractions—and are the only integers within the Figure 2 analysis sets that reproduce the paper’s estimates and Wilson intervals.

The system classified retinal images well. The second study adds useful camera evidence, but not 317 new people or an independent replication. The abstract states that “Imageability was >99% in all studies,” and a separate company page repeats “>99% imageability.” Yet Section 3.3 prints AEYE-3 at exactly 99%, with a 95% interval running down to 96.97%. The only integer fraction within its Figure 2 analysis set that reproduces that interval is 331/335, or 98.81%, which is consistent with the paper’s rounded 99%. The paper does not supply the exact underlying fractions.

This is not an accuracy takedown. It is the first example of the problem that runs through the public record on retinal AI: a claim may be supported at one rung of the care chain and then travel farther than the denominator underneath it.

Ground Truth’s AI-versus-clinician scoreboard previously highlighted a sensitivity advantage in Thailand. In the national evaluation, on the same images, Google’s deep-learning system had 91.4% sensitivity for vision-threatening diabetic retinopathy versus 84.8% for regional retina-specialist overreaders (p=.024), while specificity was effectively identical: 95.4% versus 95.5% (p=.98). We marked the patient-outcome field “No.” This investigation begins there.

The question is no longer only whether AI can read the retina. It is whether the person behind the image reaches confirmation, treatment and preserved sight.

One person, two camera-study entries

The headline cohort sum is not a unique-person count. The lineage keeps repeated participation visible without erasing the camera-specific accuracy result.

1,213 camera-study entries are at most 896 unique screened people

Three reported camera-study cohorts · AEYE-1 531 + AEYE-2 317 + AEYE-3 365 = 1,213 entriesThree reported camera-study cohortsAEYE-1 531 + AEYE-2 317 + AEYE-3 365 = 1,213entriesNoteFigure 2 shows that the 1,210 total uses 362enrolled in AEYE-3 rather than its 365screened, switching cohort stages.AEYE-1 · 531 participants · Topcon NW400; contains the 317-person AEYE-2 paired-camera subsetAEYE-1 · 531participantsTopcon NW400;contains the317-person AEYE-2paired-camera subsetAEYE-3 · 365 participants · Aurora portable camera; separately enrolled cohortAEYE-3 · 365participantsAurora portablecamera; separatelyenrolled cohortNote — paired lineageAEYE-2 reuses AEYE-1 participants; it is notan independent replication.At most 896 unique screened people · 531 + 365; at least 317 of 1,213 entries (26.1%) repeat a participantAt most 896 unique screened people531 + 365; at least 317 of 1,213 entries (26.1%)repeat a participant
Source: AEYE pivotal synthesis, §3.1, Fig. 2 and §3.3/Table 3; Ground Truth cohort reconstruction.
View chart data
1,213 camera-study entries are at most 896 unique screened people flow data
StageWhat happensNote
Three reported camera-study cohortsAEYE-1 531 + AEYE-2 317 + AEYE-3 365 = 1,213 entriesFigure 2 shows that the 1,210 total uses 362 enrolled in AEYE-3 rather than its 365 screened, switching cohort stages.
(paired lineage) AEYE-1 · 531 participantsTopcon NW400; contains the 317-person AEYE-2 paired-camera subsetAEYE-2 reuses AEYE-1 participants; it is not an independent replication.
(separate cohort) AEYE-3 · 365 participantsAurora portable camera; separately enrolled cohort
At most 896 unique screened people531 + 365; at least 317 of 1,213 entries (26.1%) repeat a participant
Read before citing
  • Repeated participation does not invalidate the camera-specific accuracy estimates. It limits the unique-person total and the independence of AEYE-2.
  • Figure 2 resolves the three-entry arithmetic gap: 531 + 317 + 362 = 1,210. The total mixes screened and enrolled stages and does not resolve the repeated-person count.

The camera-specific diagnostic matrices still reproduce

Rows are reference-center more-than-mild diabetic-retinopathy status; columns are AEYE-DS status. Counts are shown as true positive, false negative, true negative and false positive.

The camera-specific diagnostic matrices still reproduce table
Camera studyScreenedTrue positive / false negative / true negative / false positiveSensitivitySpecificityReconstructed imageability
AEYE-1 · Topcon NW40053153 / 4 / 370 / 3592.98%91.36%462/466 · 99.14%
AEYE-2 · Aurora · paired subset31734 / 3 / 233 / 1691.89%93.57%286/288 · 99.31%
AEYE-3 · Aurora · separate cohort36537 / 3 / 258 / 3392.50%88.66%331/335 · 98.81%
Source: AEYE pivotal synthesis, archived Figure 2, §3.3 and Table 3; Ground Truth diagnostic_2x2.csv and diagnostic_metrics.csv.
Read before citing
  • The paper prints exact imageability point estimates for AEYE-1 and AEYE-2, a rounded 99% for AEYE-3, and Wilson intervals, but no fractions. Ground Truth uniquely reconstructed the fractions within the Figure 2 bounds.
  • The fractions are not all-screened technical-yield counts. AEYE-2 reuses 317 AEYE-1 participants and is camera evidence, not an independent-patient replication.
  • The archived Figure 2 separately supports the three retained 2×2 matrices.

The part AI often does well

Diabetic retinopathy is a strong use case for medical AI. It is common, retinal photographs are standardized, and sight-threatening disease can be treated if people are found and cared for in time. An autonomous system can move image interpretation into a primary-care clinic, pharmacy or community program and return an answer while the patient is still there.

Across 82 studies covering 887,244 examinations and 25 regulator-approved systems, a 2025 systematic review reported pooled patient-level sensitivity of 93% and specificity of 90%. Those are high aggregate accuracy estimates. The public files do not include the patient-level 2×2 cells needed for us to reproduce the pooled results independently, and performance—especially specificity—varies substantially across settings. The review nevertheless supports the narrower conclusion that regulator-approved retinal AI can identify referable disease from retinal photographs.

The condition is that the system first has to produce an answer.

“Imageability” is not a property of software alone. It changes with the camera, the operator, the number of capture attempts, whether pupils are dilated, and what happens to an ungradable image. In a Mayo Clinic deployment, 580 of 1,052 people had AI-gradable photographs before dilation. After the protocol offered reflex dilation and repeat imaging for initially ungradable photographs, 965 ultimately had an AI-gradable set. In a Johns Hopkins deployment without reflex dilation, only 118 of 241 received a diagnostic output. A 2026 German evaluation reported 555 definitive results from 875 attempts under its no-retake, no-dilation workflow.

Those studies do not establish that one product is intrinsically more imageable than another. They show that acquisition policy is part of performance. A headline percentage that begins after unusable images have been excluded answers a different question from the proportion of all people who walked in and left with a result.

Diagnostic performance is the strongest link in the evidence chain. Technical yield and downstream care remain setting- and workflow-dependent.

A referral is not treatment

The words around screening make the care chain sound short. A camera finds disease; the patient is referred; treatment prevents blindness.

The measurable chain is longer:

image attempted → definitive result → positive result → referral recommended → referral attended → disease confirmed → treatment indicated → treatment or management documented → treatment completed → vision measured

We built that ladder into a structured dataset and searched prospective and real-world diabetic-retinopathy AI studies published from 2020 through 30 August 2026. The pathway census contains 17 studies. For comparative studies, the numeric ledger selects the AI arm; otherwise it selects the single implementation cohort. That produces 17 pathway rows with a selected baseline of 14,261 people. The baseline is usually screening attempts when reported, then arm or cohort enrollment. One source-defined exception is explicit: Google Thailand uses its 7,651-person analysis cohort after 7,940 people were screened for inclusion.

We extracted every patient count each study reported. Because studies stop reporting at different stages, no single group of 14,261 people can be followed from the first row to the last. Read each line independently: it is the minimum documented count among studies reporting that stage, not a conversion rate.

Across changing contributor sets, the ledger records 3,049 definitive AI results, 918 cases with no definitive result or an ungradable image, 3,773 people eligible for referral, 630 recorded attendances, 477 confirmatory examinations and 90 confirmed disease cases. Eligible and attempted totals can exceed 14,261 because studies use different enrollment, screening and analysis denominators. The matching imageable and definitive-result totals come from the same eight studies, although the two concepts remain distinct. The pathway ledger below preserves the full reconciliation and every study row.

How many patients needed treatment? Only 10, across two studies, were explicitly counted as needing it. How many received treatment? Thirty-three patients across five studies had a named treatment or a clearly documented first-treatment event. A broader count reaches 56 across six studies by adding 23 Maine patients described only as receiving unspecified “appropriate treatments.” The 33 are included in the 56; the 10 come from a different study set and overlap only partly. These are overlapping evidence layers, not steps in one funnel.

Across the full 14,261-person census, those counts are minimum observed shares of 0.07%, 0.23% and 0.39%. Within the studies contributing each row, they equal 10/85, 33/289 and 56/334 of recorded attendees. Because each row uses a different study set, those percentages are not sequential conversion rates. Many attendees appropriately did not need treatment; only EyeArt directly observed indication followed by treatment, in three of six patients.

The Maine group also appears in the disease-confirmed count. Table 1 reports 23 of 45 follow-up patients as diabetic-retinopathy positive; the text says 51% “were found to have [diabetic retinopathy] and received appropriate treatments.” We therefore carry one reconstructed 23-person group in both evidence layers—not a separately observed 23→23 transition. The source does not define modality, timing, completion or whether “appropriate treatments” included observation.

The missing data are concentrated. Google Thailand supplies 7,651 of the 14,261 selected patients, 53.6%. Rows with no attendance count supply 8,822 patients, 61.9% of the census; rows with no treatment or management count supply 10,267, 72.0%. These figures describe the selected public evidence—not an average program or representative patient population.

Some apparent transitions are not nested:

  • Australia’s 74 AI referral flags and 28 clinical referrals use different criteria.
  • EyeArt’s 92 referrals include all 53 inconclusive outputs, so they are not a subset of its 127 definitive AI results.
  • Excluding those rows leaves a valid definitive-result-to-referral link of 462/2,219 across five studies, 20.8%.

Other rows enter the ledger at different points:

  • Google’s national Thailand paper reports 7,940 people screened for inclusion and a 7,651-person analysis cohort. The selected baseline is that analysis cohort. Its 2,412 composite referrals make up 63.9% of the 3,773 referral-eligible column sum, but include diabetic retinopathy, diabetic macular edema, ungradable images and poor visual acuity. They are not an AI-positive count.
  • RAIDERS contributes a 136-person post-positive randomized arm. DeepDR-LLM begins with 144 already referral-selected patients.
  • Jordan contributes 402 screened people to the census, but its downstream percentage-derived counts conflict. They remain marked with question marks and excluded from stage totals and links.

We used strict definitions. Advice to seek care is not attendance. Attendance is not confirmed disease. “Treatment required” is not treatment delivered. A first laser or injection is evidence of initiation, not proof that the intended course was completed. Same-day specialist examinations among patients already inside an eye hospital were not recoded as community referral uptake.

No study in the 17-study census documents completion of the indicated treatment plan or reports a denominator for longitudinal visual outcome. That does not mean nobody completed treatment or preserved sight. It means the public studies we found cannot tell us how often either happened.

The 17-study patient-count ledger

Every cell is an extracted canonical-stage patient count from one selected AI or single-pathway row. Both percentage columns are documented-minimum ratios; valid links use nested counts within rows.

What 17 selected pathway rows actually counted

14,261 patients across 17 selected pathway rows. Use an explicit source-defined analysis cohort where declared; otherwise use screening attempts when reported, then selected arm or cohort enrollment. Both percentage columns are minimum documented shares, not incidence or complete-ascertainment estimates.

Treatment evidence, in three overlapping layers

Three overlapping layers of reported treatment evidence
Evidence layerPeopleStudiesRecorded minimum in all 14,261Count ÷ recorded attendees in the same studiesWhat the studies documented
Indication explicitly reported102/170.07%10/85 · 11.8%Two studies separately counted who was judged to need treatment.
Specified treatment event335/170.23%33/289 · 11.4%Named modality or clearly documented first-treatment event; contained within the inclusive 56.
Inclusive treatment / management event566/170.39%56/334 · 16.8%Adds Maine’s same 23 people described only as receiving unspecified ‘appropriate treatments.’
Do not connect these rows Only EyeArt independently observes indication followed by treatment: 3 of 6. The three rows above use different contributor sets and must not be added or connected as a funnel.
Where the missing data sit Google Thailand supplies 7,651/14,261 patients (53.6%). Rows with no attendance count supply 8,822 (61.9%); rows with no treatment/management count supply 10,267 (72.0%). Missingness is concentrated, not random.

Counts and percentages

Column sums—not a common cohort. The contributor set and reporting-row baseline change by stage.

Stage counts and percentages across the 17-study pathway census
StageDocumented nn / full 14,261Studies reportingReporting-row baselinesDocumented n / reporting-row baselines
Eligible / enrolled14,646Not comparable16/1713,941Not comparable
Screening attempted14,270Not comparable15/1713,981Not comparable
Definitive AI result3,04921.4%8/173,96776.9%
No definitive result or ungradable9186.44%8/173,96723.1%
Study-defined AI referral flag9776.85%10/174,99519.6%
Study-defined referral eligible3,77326.5%15/1713,36528.2%
Follow-up status ascertained1,0947.67%11/174,91922.2%
Recorded referral attendance6304.42%13/175,43911.6%
Confirmatory examination4773.34%9/174,32611.0%
Disease confirmed†900.63%3/171,2087.45%
Treatment indication independently reported100.07%2/174362.29%
Treatment or management event documented560.39%6/173,9941.40%
Treatment course completedUnknown0/17Unknown
Longitudinal vision outcome measuredUnknown0/17Unknown

† Maine supplies the same reconstructed 23-person group to disease confirmed and the inclusive treatment/management count; this is one reported group, not two consecutive events.

Valid within-row links

Explore every extracted patient-count row (17 studies)

Every extracted canonical-stage patient count

Symbol key: — not reported · ? conflicting percentage-only reports without an exact n · † source count shares a reported group across stages or lacks a separately enumerated indication; see its row note.

Screening and triage
Screening and triage patient counts in all 17 pathway studies
Study pathway rowEligibleAttemptedImageableDefinitive resultNo definitive resultReferral flagReferral eligible
STATUS · AIUnited States · baseline 2,243 (screening attempted)2,243 attempts → 1,459 definitive AI results + 784 without a definitive result or ungradable → 279 positive → 99 internal visits → 9 first-visit treatments.2,2432,2431,4591,459784279279
Thailand platform · AIThailand · baseline 708 (screening attempted)201 AI flags → 129 over-reader-retained referrals → 115 attendances → 48 with vision-threatening diabetic retinopathy → 18 treated.708708201129
MogaIndia · baseline 343 (screening attempted)34334364
Rural MaineUnited States · baseline 320 (screening attempted)The same 23-person diabetic-retinopathy-positive group is described as receiving unspecified ‘appropriate treatments.’ Ground Truth reconstructed 23 by linking the paper's 51%-of-45 sentence to Table 1's 23 diabetic-retinopathy-positive patients; modality, timing and whether this included observation are unreported.3206161
Boothgarh · home AIIndia · baseline 200 (screening attempted)The 86 referrals included 36 diabetic-retinopathy-positive and 50 ungradable screens in the supplement; the linked gradability result is eye-level.20020086
EyeArt follow-up · AIUnited States · baseline 180 (screening attempted)The 92 referrals include all 53 inconclusive outputs, so they are not a subset of the 127 definitive AI results.185180127127533992
AEYE real-worldIsrael · baseline 256 (screening attempted)256256245245117676
Northern OntarioCanada · baseline 202 (screening attempted)202202189189134242
RAIDERS · AIRwanda · baseline 136 (eligible)Nested AI arm after 827 screened → 823 analyzed → 275 positive/randomized → 136 assigned AI.136136
ACCESS · AIUnited States · baseline 81 (screening attempted)Study lineage: 170 candidates → 164 randomized → 81 assigned to the selected AI row.8181818102525
Northern IndiaIndia · baseline 390 (screening attempted)390390159
AustraliaAustralia · baseline 236 (screening attempted)The 74 AI-positive outputs and 28 clinical referrals use different definitions; they are not a direct link.45623623223247428
DeepDR-LLM · AIChina · baseline 144 (eligible)The 144-person baseline is an already referral-selected RDR cohort, not an all-screened population.144144
Google ThailandThailand · baseline 7,651 (analysis cohort)The source reports 7,940 screened for inclusion and 7,651 eligible for analysis. The selected census baseline is the 7,651-person analysis cohort. Its 2,412 composite referrals combine diabetic retinopathy, diabetic macular edema, ungradable images or poor visual acuity; they are not an AI-positive count.7,6517,9402,412
BelizeBelize · baseline 275 (screening attempted)639275245245304040
B-PRODUCTIVE · AIBangladesh · baseline 494 (screening attempted)Parent trial: 993 eligible → 494 AI / 499 control. Side branch: 140 positive + 23 insufficient → 163 same-day specialist exams at eye clinics.49449447147123140
Jordan pharmacyJordan · baseline 402 (screening attempted)Percentage-derived counts are 68 positive and 12 ungradable; reported referral completion implies 51/63, but the published chain conflicts and is quarantined.51840268??
Column sum—not a common cohort14,64614,2703,0493,0499189773,773
Studies reporting16/1715/178/178/178/1710/1715/17
Downstream care
Downstream care patient counts in all 17 pathway studies
Study pathway rowFollow-up knownAttendedConfirm examConfirmedIndication reportedTreatment / managementCompletedVision outcome
STATUS · AIUnited States · baseline 2,243 (screening attempted)27999999
Thailand platform · AIThailand · baseline 708 (screening attempted)1291151154818
MogaIndia · baseline 343 (screening attempted)2891
Rural MaineUnited States · baseline 320 (screening attempted)† The same reconstructed 23 people supply both confirmed disease and unspecified treatment/management; the paper does not establish modality, timing, or that active treatment rather than observation was intended.454523†23†
Boothgarh · home AIIndia · baseline 200 (screening attempted)† Two people had recorded diabetic-retinopathy-treatment status, but treatment indication was not separately enumerated.15152†
EyeArt follow-up · AIUnited States · baseline 180 (screening attempted)9251511963
AEYE real-worldIsrael · baseline 256 (screening attempted)7634344
Northern OntarioCanada · baseline 202 (screening attempted)423232
RAIDERS · AIRwanda · baseline 136 (eligible)1367070
ACCESS · AIUnited States · baseline 81 (screening attempted)251616
Northern IndiaIndia · baseline 390 (screening attempted)12823
AustraliaAustralia · baseline 236 (screening attempted)159
DeepDR-LLM · AIChina · baseline 144 (eligible)144112
Google ThailandThailand · baseline 7,651 (analysis cohort)
BelizeBelize · baseline 275 (screening attempted)
B-PRODUCTIVE · AIBangladesh · baseline 494 (screening attempted)
Jordan pharmacyJordan · baseline 402 (screening attempted)??
Column sum—not a common cohort1,094630477901056
Studies reporting11/1713/179/173/172/176/170/170/17
Every displayed value is an extracted patient count—not a yes/no endpoint flag. Column sums are not a common cohort. Percentages use either the 14,261-patient selected-pathway census, reporting-row baselines, or valid nested counts for one labelled link.
Source: Ground Truth numeric pathway census; outputs/pathway_numeric_census.csv, pathway_numeric_stage_totals.csv and pathway_numeric_links.csv; public evidence searched through 30 Aug 2026.
Read before citing
  • This is one selected pathway row per study, not a homogeneous screened population. RAIDERS contributes a 136-person post-positive AI arm and DeepDR-LLM an already referral-selected 144-person cohort.
  • Both percentage columns are minimum documented shares, not incidence or care-completion estimates. An omitted stage remains unknown, not zero.
  • Stage totals use different contributor sets and must not be connected as attrition. The 10, 33 and 56 treatment counts are overlapping evidence layers, not consecutive gates.
  • Each link bar uses only rows with valid nested counts at both ends. EyeArt and Australia are excluded from the definitive-result-to-referral link because their referral counts are not subsets of definitive AI results.
  • Jordan contributes 402 screened people to the baseline, but its downstream percentage-derived counts conflict; they are shown with question marks and excluded from canonical stage totals and links.
  • No study reported completed treatment or a denominator for longitudinal vision outcome. Those percentages are unknown, not 0%.

Sometimes the missing denominator nearly doubles the rate

Even “attended referral” is not one number when researchers cannot observe every patient.

In an Australian implementation study, 28 people were referred, 15 could be contacted three months later, and nine said they had attended. That is 32.1% of everyone referred or 60% of the people reached. The second figure describes attendance among people whose status is known; the first is the conservative recorded minimum. Neither reveals what happened to the 13 people who could not be reached.

The same distinction matters in Moga, India. Of 64 people referred, 28 were contacted one month later and nine said they had followed the advice and visited an ophthalmologist: 14.1% of everyone referred, or 32.1% of those contacted. We downgrade that to a reported referral attempt because five of the nine went to facilities without eye care.

The paper then lists 10 dispositions for nine people—one optical-coherence-tomography referral, those five visits, one eye-drop treatment, one recommendation for an anti–vascular endothelial growth factor eye injection, one laser treatment and one follow-up visit. The categories overlap or the report contains a counting error. We retain nine as the source-reported referral count, not nine confirmed examinations. Only the laser is coded as delivered diabetic-retinopathy treatment; the eye drops are not identified as diabetic-retinopathy directed.

In the US STATUS program, 99 of 279 AI-positive patients had a directly observed Stanford ophthalmology visit within 90 days. The paper’s 69.2% any-provider estimate adds those internal visits, 35.5%, to 33.7% attributed to community visits. Ground Truth translates 33.7% of 279 to approximately 94 people, but the paper does not print 94 or show the exact extrapolation.

Outside follow-up was estimated from a phone sample: 152 people were called, 93 were reached and 68 of 93 reported follow-up. The paper labels 93/152 as 63.4%; the arithmetic is 61.2%. Our ledger retains the directly observed 99.

STATUS also had a nested AI–human salvage branch. Of 784 AI-ungradable cases, 664 entered human review; 561 were human-negative, 20 human-positive and 83 remained ungradable. Twelve of the paper’s own referrable denominator of 103 had an internal visit. The figure does not account for the other 120 AI-ungradable cases, so we show the hybrid branch as context instead of adding it as an independent cohort.

Internal records cannot rule out outside care. A patient absent from that record may have gone nowhere or may have gone elsewhere. “Not observed here” is not the same as “did not attend.”

This sounds like bookkeeping because it is bookkeeping. It is also the difference between a care gap and a data gap.

The pooled effect survives. Its simplicity does not

A 2026 systematic review and meta-analysis identified six comparative studies of AI-assisted diabetic-retinopathy pathways. We reconstructed all six 2×2 tables and reproduced the published result to the reported precision.

Across the six, recorded referral attendance was higher in the AI-enabled pathway: pooled risk ratio 1.89, with a 95% confidence interval from 1.18 to 3.03. The pooled absolute difference was 23.9 percentage points, from 12.8 to 35.1.

This is the strongest comparative evidence we found that AI-enabled pathways are associated with higher recorded referral attendance.

It is not a clean estimate of what the algorithm itself caused.

The table below reproduces the review’s published-synthesis extraction:

Referral attendance in six comparative studies
Comparative study AI-pathway attendance Comparator attendance Risk ratio
ACCESS, United States 16/25, 64.0% 18/83, 21.7% 2.95
DeepDR-LLM, China 112/144, 77.8% 90/154, 58.4% 1.33
EyeArt follow-up, United States 51/92, 55.4% 182/974, 18.7% 2.97
RAIDERS, Rwanda 70/136, 51.5% 55/139, 39.6% 1.30
STATUS, United States 99/279, 35.5% 14/117, 12.0% 2.97
Thailand digital platform 115/129, 89.1% 124/175, 70.9% 1.26

STATUS also reported 12 internal visits among 103 referral-eligible patients in its nested AI–human salvage branch. The table follows the review’s AI-versus-historical-control comparison; the hybrid branch is not omitted from the evidence record or added as an independent cohort.

For ACCESS, the review uses 18/83 in the control arm. The primary analysis uses 18/82 after excluding one control participant enrolled in another screening study; our primary-endpoint sensitivity analysis uses 18/82.

ACCESS is also stage-asymmetric by design. The AI numerator is 16 of 25 participants with a disease-present result who completed a follow-up eye-care visit. The control numerator is 18 participants who completed an initial diabetic-eye examination after routine referral; all 18 were found not to have disease. That is a defensible end-to-end pathway contrast, but not a like-for-like referral-uptake comparison among screen-positive patients. A stage-matched control risk ratio cannot be recovered because the control arm did not identify an equivalent disease-positive denominator.

The studies disagree sharply. I², a measure of between-study inconsistency, was 91.9%. The 95% prediction interval—our estimate of what a new similar setting might show—ran from a risk ratio of 0.53 to 6.78, spanning a possible decrease through a very large increase.

We also repeated the analysis while omitting one study at a time. With the review’s published rows, one of six confidence intervals crossed 1; after correcting the ACCESS control denominator and substituting Thailand’s closer stage comparison, three of six did.

The comparison arms differ too. In ACCESS, the control group received a routine referral recommendation while the AI group received a point-of-care result. In RAIDERS, the randomized contrast was an immediate AI result against delayed human grading. Other studies add counselling, scheduling, reminders or new referral rules. Our coding found at least two workflow differences in every comparison and as many as four; unreported components may raise that count.

Remove ACCESS and EyeArt—the two comparisons whose controls received routine or universal referral advice—and align Thailand to the same care stage in both arms. The remaining DeepDR-LLM, RAIDERS, STATUS and Thailand rows produce a descriptive pooled risk ratio of 1.47, with a confidence interval from 0.80 to 2.71. That restriction removes two of the three largest individual risk ratios. The two randomized trials alone give a risk ratio of 1.90 with a 95% confidence interval from 0.01 to 341.83, far too wide to guide a stable conclusion.

The pooled average favors AI-enabled pathways. Whether that result will travel to a new setting is unresolved.

At least two rows do not compare the same clinical transition in both arms. ACCESS is one. Thailand is the clearest numerical mismatch.

In the implementation study, 115 of 129 AI-pathway true-positive referrals attended tertiary care. The review’s narrative compares that 89.1% with the primary paper’s stage-matched 17 of 22, or 77.3%, and reports p=.158. Yet its synthesis table uses 124 of 175, or 70.9%, for the manual period; that 124 appears to combine 116 intermediate confirmation visits and eight direct tertiary visits. The periods are observational and not directly comparable, but the review still pooled the less aligned of its own two manual-period numbers.

Correcting only that row barely changes the pooled point estimate: risk ratio 1.87, with a 95% confidence interval from 1.14 to 3.06. The result survives. The interpretation changes. The original rows do not all estimate the same clinical transition, and the corrected prediction interval remains wide, 0.49 to 7.07.

The six AI denominators in the review’s main synthesis table sum to 805. Its supplementary summary-of-findings table instead says “AI-assisted screening: 889,” 84 more. The manual-arm denominators reconcile exactly at 1,642, but neither the main paper nor its two supplements explains the extra 84 on the AI side.

That internal discrepancy remains unresolved in the public report. Neither 805 nor 889 is a shared screening denominator; both describe study-defined groups entering different comparisons.

On average, AI-enabled pathways recorded more referral attendance, but they tested whole workflows, not the algorithm alone. The review summarized the absolute difference as roughly one extra referral completion per four eligible people. That is not a screening-wide number needed to treat: eligibility varied across studies, and our prediction interval for the absolute difference ran from −1.6 to +49.5 percentage points, crossing zero.

A positive mean, a wide range of plausible settings

The forest separates confidence intervals from prediction intervals. The matrix keeps the surrounding workflow bundle visible.

AI-enabled pathways raised recorded attendance on average—not reliably everywhere

0.51248ACCESSACCESS: 2.95 (95% confidence interval 1.78 to 4.88; 16/25 vs 18/83)ACCESS: 2.95 (95% confidence interval 1.78 to 4.88; 16/25 vs 18/83)2.95 [1.78, 4.88]DeepDR-LLMDeepDR-LLM: 1.33 (95% confidence interval 1.13 to 1.56; 112/144 vs 90/154)DeepDR-LLM: 1.33 (95% confidence interval 1.13 to 1.56; 112/144 vs 90/154)1.33 [1.13, 1.56]EyeArt follow-upEyeArt follow-up: 2.97 (95% confidence interval 2.37 to 3.72; 51/92 vs 182/974)EyeArt follow-up: 2.97 (95% confidence interval 2.37 to 3.72; 51/92 vs 182/974)2.97 [2.37, 3.72]RAIDERSRAIDERS: 1.30 (95% confidence interval 1.00 to 1.69; 70/136 vs 55/139)RAIDERS: 1.30 (95% confidence interval 1.00 to 1.69; 70/136 vs 55/139)1.30 [1.00, 1.69]STATUSSTATUS: 2.97 (95% confidence interval 1.77 to 4.97; 99/279 vs 14/117)STATUS: 2.97 (95% confidence interval 1.77 to 4.97; 99/279 vs 14/117)2.97 [1.77, 4.97]Thailand · publishedThailand · published: 1.26 (95% confidence interval 1.12 to 1.41; 115/129 vs 124/175)Thailand · published: 1.26 (95% confidence interval 1.12 to 1.41; 115/129 vs 124/175)1.26 [1.12, 1.41]Thailand · same-stageThailand · same-stage: 1.15 (95% confidence interval 0.91 to 1.46; 115/129 vs 17/22) — interval includes 1Thailand · same-stage: 1.15 (95% confidence interval 0.91 to 1.46; 115/129 vs 17/22) — interval includes 11.15 [0.91, 1.46]Published pooledPublished pooled: 1.89 (95% confidence interval, adjusted for six studies 1.18 to 3.03; 6 studies)Published pooled: 1.89 (95% confidence interval, adjusted for six studies 1.18 to 3.03; 6 studies)1.89 [1.18, 3.03]Same-stage pooledSame-stage pooled: 1.87 (95% confidence interval, adjusted for six studies 1.14 to 3.06; Thailand alternate)Same-stage pooled: 1.87 (95% confidence interval, adjusted for six studies 1.14 to 3.06; Thailand alternate)1.87 [1.14, 3.06]Published new-setting rangePublished new-setting range: 1.89 (95% prediction interval 0.53 to 6.78; I² 91.9%) — interval includes 1Published new-setting range: 1.89 (95% prediction interval 0.53 to 6.78; I² 91.9%) — interval includes 11.89 [0.53, 6.78]Same-stage new-setting rangeSame-stage new-setting range: 1.87 (95% prediction interval 0.49 to 7.07; Thailand alternate) — interval includes 1Same-stage new-setting range: 1.87 (95% prediction interval 0.49 to 7.07; Thailand alternate) — interval includes 11.87 [0.49, 7.07]Risk ratio · comparator higher ← 1 → AI-enabled pathway higher
Circles show individual study rows; diamonds show pooled estimates; amber bands show 95% prediction intervals for a new setting.
Source: Ground Truth reconstruction of six published 2×2 referral tables; outputs/comparative_effects.csv and findings.json.
View chart data
AI-enabled pathways raised recorded attendance on average—not reliably everywhere data table
ComparisonRisk ratioLimitsInterval typeRecorded attendanceInterval vs 1Design note
ACCESS2.951.78 to 4.8895% confidence interval16/25 vs 18/83excludes 1Randomized controlled trial; low risk of bias; stage-asymmetric transition
DeepDR-LLM1.331.13 to 1.5695% confidence interval112/144 vs 90/154excludes 1Sequential prospective comparison; moderate risk of bias
EyeArt follow-up2.972.37 to 3.7295% confidence interval51/92 vs 182/974excludes 1Prospective cohort with historical comparator; serious risk of bias
RAIDERS1.301.00 to 1.6995% confidence interval70/136 vs 55/139excludes 1Randomized controlled trial; low risk of bias
STATUS2.971.77 to 4.9795% confidence interval99/279 vs 14/117excludes 1Historical workflow comparison; serious risk of bias
Thailand · published1.261.12 to 1.4195% confidence interval115/129 vs 124/175excludes 1Alternating implementation periods; serious risk of bias
Thailand · same-stage1.150.91 to 1.4695% confidence interval115/129 vs 17/22includes 1Alternative extraction of the same study; not a seventh comparison
Published pooled1.891.18 to 3.0395% confidence interval, adjusted for six studies6 studiesexcludes 1Random effects; between-study inconsistency (I²) 91.9%
Same-stage pooled1.871.14 to 3.0695% confidence interval, adjusted for six studiesThailand alternateexcludes 1Sensitivity analysis; I² 90.9%
Published new-setting range1.890.53 to 6.7895% prediction intervalI² 91.9%includes 1Expected range for a new setting; not a confidence interval
Same-stage new-setting range1.870.49 to 7.0795% prediction intervalThailand alternateincludes 1Prediction interval after stage-aligning the Thailand row
Read before citing
  • Both the published and same-stage pooled results are low-certainty, highly heterogeneous and pathway-specific. Their 95% prediction intervals include lower attendance in a new setting.
  • The same-stage Thailand point is an alternative extraction of the same study, not a seventh independent comparison.
  • ACCESS is stage-asymmetric by design: 16/25 is follow-up among AI disease-present patients, while 18/83 is an initial eye examination among the entire routinely referred control arm. It is not a like-for-like uptake comparison among screen-positive patients.
  • The six rows use study-defined denominators and mixed designs. They estimate bundled pathways, not the classifier alone. The main-table AI denominators sum to 805; Supplementary Table 7 prints 889, an unexplained difference of 84.
  • In the accessible table, confidence-interval limits are the lower and upper bounds around each estimate. Prediction intervals describe the wider range a new similar setting might show.

Every comparison changed more than the classifier

Every comparison changed more than the classifier workflow matrix
ComparisonInstant resultSame-day optionCounsellingSchedulingRemindersNavigation / transportChanged referral rule
ACCESS2 added; routine or universal referraladded in AI pathwaynot documented in the AI pathwaydocumented in both pathwaysnot documented in the AI pathwaynot documented in the AI pathwaynot documented in the AI pathwayadded in AI pathway
DeepDR-LLM3 added; manual grading workflowadded in AI pathwaynot documented in the AI pathwayadded in AI pathwaynot documented in the AI pathwaynot documented in the AI pathwaynot documented in the AI pathwayadded in AI pathway
EyeArt follow-up4 added; routine or universal referraladded in AI pathwaynot documented in the AI pathwayadded in AI pathwayadded in AI pathwaynot documented in the AI pathwaynot documented in the AI pathwayadded in AI pathway
RAIDERS2 added; manual grading workflowadded in AI pathwayadded in AI pathwaydocumented in both pathwaysnot documented in the AI pathwaynot documented in the AI pathwaydocumented in both pathwaysnot documented in the AI pathway
STATUS4 added; historical workflowadded in AI pathwaynot documented in the AI pathwayadded in AI pathwayadded in AI pathwaynot documented in the AI pathwaynot documented in the AI pathwayadded in AI pathway
Thailand3 added; manual grading workflowadded in AI pathwaynot documented in the AI pathwaydocumented in both pathwaysdocumented in both pathwaysadded in AI pathwaynot documented in the AI pathwayadded in AI pathway
● added in AI pathway   ○ documented in both pathways   — not documented in the AI pathway
Source: Ground Truth extraction; outputs/workflow_components.csv and primary-study Methods/Results.
Read before citing
  • A filled dot means the component was documented in the AI pathway and not documented as present in its comparator. It does not identify which component caused the attendance difference.
  • Open dots keep shared counselling, scheduling or transport visible so those co-interventions are not silently credited to AI.

No synthetic “average patient journey” can be made by splicing the best number from one study onto the next number from another. Three individual programs show why.

AEYE in an Israeli endocrinology clinic. In the 2026 real-world study, 256 people had an exam attempted, 245 received a definitive result and 76 were positive. The originating clinic documented 34 confirmatory examinations. Four of those 34 were judged to require treatment. The other 42 had no recorded internal confirmation but cannot be classified as having received no confirmatory care, because outside care was not captured. The paper does not report whether the four started or completed treatment, or what happened to vision.

Home AI screening in Boothgarh, India. In a three-arm pragmatic study, 200 people were enrolled in the home AI arm, 86 were referred and 15 reported by phone one month later that they had attended. Two participants had recorded diabetic-retinopathy treatment: one received laser, and one received laser plus an anti–vascular endothelial growth factor eye injection (supplementary Table 6). The paper does not report the intended treatment course, course completion or visual outcomes.

A linked publication from the same setting reported image gradability at eye level—224 of 362 analyzed eyes in the home/community AI arm—and that figure cannot be inserted as a patient-level step. The main adoption article and supplement also conflict on arm-level phone ascertainment, so we report attendance against the full referral group instead of selecting the more favorable contacted denominator.

A digital platform in Thailand. Here the chain extends farthest. Of 708 people screened in the AI period, 63 entered a separate visual-acuity referral branch. Among the remaining 645, AI flagged 201; specialist overread retained 129; 115 attended tertiary care; 48 had vision-threatening diabetic retinopathy; and 18 received laser or an eye injection.

The side branches matter. Of the 63 visual-acuity referrals, 52 attended. Overread rejected 72 AI positives, although 17 still presented, and found five false negatives, four of whom attended. The 67 ungradable attenders included two cataract treatments, which we do not count as diabetic-retinopathy treatment. No completed treatment course or vision outcome was reported.

The Thailand study documents high recorded attendance within a redesigned pathway; it does not isolate an AI effect or establish superiority over the manual period. The program combined AI, specialist overread, reminders and active referral cancellation. Credit belongs to the pathway.

Three programs, three separate missing links

The cascades stay separate because their gates, capture domains and endpoints differ. Side branches remain visible rather than being converted into attrition.

AEYE real-world clinic: outside follow-up breaks the chain

Exam attempted · 256 peopleExam attempted · 256 peopleDefinitive result · 245 · 95.7% of attemptsDefinitive result · 24595.7% of attemptsPositive result · 76 · 31.0% of definitive results; 29.7% of attemptsPositive result · 7631.0% of definitive results; 29.7% of attemptsInternal confirmation · 34 · 44.7% of positive resultsInternal confirmation· 3444.7% of positiveresultsNo internal confirmation record · 42 · Cannot be classified as no confirmatory careNo internalconfirmation record ·42Cannot be classifiedas no confirmatorycareTreatment judged necessary · 4 · 11.8% of internally examined patientsTreatment judged necessary · 411.8% of internally examined patientsInitiation · completion · vision · Not reportedInitiation · completion · visionNot reported
Source: AEYE 2026 Israeli endocrinology-clinic study, PMID 42373309, Methods and Results.
View chart data
AEYE real-world clinic: outside follow-up breaks the chain flow data
StageWhat happensNote
Exam attempted · 256 people
Definitive result · 24595.7% of attempts
Positive result · 7631.0% of definitive results; 29.7% of attempts
(observed internally) Internal confirmation · 3444.7% of positive results
(outside care unobserved) No internal confirmation record · 42Cannot be classified as no confirmatory care
Treatment judged necessary · 411.8% of internally examined patients
Initiation · completion · visionNot reported
Read before citing
  • The 42 people without an internal confirmation record may include outside care. They are unobserved, not verified nonattenders.
  • Treatment judged necessary is not evidence that treatment began or was completed.

Boothgarh home AI: treatment status recorded for two; completion unknown

Home AI arm · 200 peopleHome AI arm · 200 peopleReferred · 86 · 43.0% of the armReferred · 8643.0% of the armRecorded attendance at one month · 15 · 17.4% of referralsRecorded attendance at one month · 1517.4% of referralsNoteThe remaining 71 have no recorded visit; thestudy does not verify all as nonattenders.Participants with recorded diabetic-retinopathy treatment · 2 · One retinal laser; one retinal laser plus a drug injection into the eye (anti-VEGF)Participants with recordeddiabetic-retinopathy treatment · 2One retinal laser; one retinal laser plus a druginjection into the eye (anti-VEGF)NoteAmong 11 ungradable attenders, the supplementlists four as 'cat surgery,' three as 'cat Sxappointment' and four as no treatment.Completed cataract surgery is not clearlyestablished; these categories are not countedas diabetic-retinopathy treatment.Completed diabetic-retinopathy episode · vision outcome · Not reportedCompleted diabetic-retinopathy episode ·vision outcomeNot reported
Source: Boothgarh pragmatic study, Results and supplement Tables 6–7; Ground Truth conservative recorded-minimum extraction.
View chart data
Boothgarh home AI: treatment status recorded for two; completion unknown flow data
StageWhat happensNote
Home AI arm · 200 people
Referred · 8643.0% of the arm
Recorded attendance at one month · 1517.4% of referralsThe remaining 71 have no recorded visit; the study does not verify all as nonattenders.
Participants with recorded diabetic-retinopathy treatment · 2One retinal laser; one retinal laser plus a drug injection into the eye (anti-VEGF)Among 11 ungradable attenders, the supplement lists four as 'cat surgery,' three as 'cat Sx appointment' and four as no treatment. Completed cataract surgery is not clearly established; these categories are not counted as diabetic-retinopathy treatment.
Completed diabetic-retinopathy episode · vision outcomeNot reported
Read before citing
  • Arm-level phone-ascertainment totals conflict between the main text and supplement, so the chart does not calculate attendance only among people reached.
  • The linked 224/362 gradability count is eye-level and is not inserted into this patient-level funnel.

Thailand: the longest reported chain still has side branches

Screened in AI period · 708Screened in AI period · 708Separate visual-acuity referral · 63 · 52 later presented at tertiary careSeparate visual-acuityreferral · 6352 later presented attertiary careEntered AI-image comparison · 645 · The 201 AI positives use this denominator—not 708Entered AI-imagecomparison · 645The 201 AI positivesuse thisdenominator—not 708AI referral-positive · 201 · 31.2% of 645AI referral-positive · 20131.2% of 645NoteOf 444 AI negatives, overread identified fivefalse negatives; four attended after contact.Over-reader agreed · 129 · 64.2% of AI positivesOver-reader agreed ·12964.2% of AI positivesOver-reader disagreed · 72 · 17 still presentedOver-reader disagreed· 7217 still presentedNote — false-positive branch49 were reached to reverse referral (2attended); 23 were not reached (15 attended).True-positive referrals attending tertiary care · 115 · 89.1% of 129True-positive referrals attending tertiarycare · 11589.1% of 129Vision-threatening diabetic retinopathy confirmed · 48 · 45 diabetic macular edema; 3 severe retinopathyVision-threatening diabetic retinopathyconfirmed · 4845 diabetic macular edema; 3 severe retinopathyRetinal laser or eye injection received · 18 · 25 monitored; 5 referred onwardRetinal laser or eye injection received · 1825 monitored; 5 referred onwardCompleted treatment episode · vision outcome · Not reportedCompleted treatment episode · vision outcomeNot reported
Source: Thailand digital-platform study, Fig. 1, Tables 1–3 and Results; Ground Truth branch reconstruction.
View chart data
Thailand: the longest reported chain still has side branches flow data
StageWhat happensNote
Screened in AI period · 708
(VA ≤20/70) Separate visual-acuity referral · 6352 later presented at tertiary care
(AI workflow denominator) Entered AI-image comparison · 645The 201 AI positives use this denominator—not 708
AI referral-positive · 20131.2% of 645Of 444 AI negatives, overread identified five false negatives; four attended after contact.
(true-positive referrals) Over-reader agreed · 12964.2% of AI positives
(false-positive branch) Over-reader disagreed · 7217 still presented49 were reached to reverse referral (2 attended); 23 were not reached (15 attended).
True-positive referrals attending tertiary care · 11589.1% of 129
Vision-threatening diabetic retinopathy confirmed · 4845 diabetic macular edema; 3 severe retinopathy
Retinal laser or eye injection received · 1825 monitored; 5 referred onward
Completed treatment episode · vision outcomeNot reported
Read before citing
  • The main spine follows over-reader-agreed true-positive referrals. It does not absorb the 52 VA-branch attenders, 17 false-positive presentations or four false-negative attenders into the 115 denominator.
  • A separate 67-person ungradable branch included two cataract treatments; those are not diabetic-retinopathy treatment events.
  • The pathway combined AI, specialist overread, reminders and referral reversal. The chart cannot isolate an algorithm-only effect.

A million screens, outcomes uncounted

The language of deployment moves more freely than the evidence.

In March, Google reported more than one million screenings through clinical partnerships across India, Thailand and Australia. The company says a diagnosis can arrive “in as little as two minutes” and “could potentially save their sight.” The sentence is appropriately hedged. Ground Truth did not independently verify the scale, and the page does not say whether each screening represents a unique person. The linked public studies do not provide one program-wide count moving from result to confirmation, completed treatment and vision.

Digital Diagnostics uses “Prevent blindness through early disease detection” as product language for LumineticsCore. The linked STATUS implementation study documents nine patients who “received treatment … at the first follow-up encounter.” That is further down the chain than diagnostic accuracy alone. It is still not a completed treatment plan or a longitudinal vision outcome.

Orbis says patients given real-time AI diagnoses “are more likely to seek treatment.” The underlying RAIDERS randomized trial measured presentation for referral services within 30 days; it did not measure treatment initiation.

Other current scale statements have the same structural limit. Eyenuk has reported more than 230,000 patients screened.

Remidio uses the same 16-million figure for different stages: its product page displays “16M+ Patients Screened,” while an NITI Frontier Tech profile submitted under the Frontier Voices programme describes “16M+ people” risk-triaged and separately lists “1.2M+ retinal scans” analyzed. The profile says those scans generated epidemiological insights, including glaucoma-prevalence estimates.

India’s MadhuNetrAI announcement says 7,100 patients are “benefiting.” None of those public numbers, as currently linked, supplies a complete program-wide treatment and vision denominator.

These statements may each be accurate at their own stage, but they are not interchangeable. Screening volume is a screening-volume claim. Referral attendance is a referral claim. Treatment initiation is a treatment claim. Preserved sight is an outcome claim.

Each requires its own denominator.

The fair case for deploying before the final outcome

A screening program does not need to wait for a definitive population-level blindness trial before using an accurate, regulated test.

The clinical logic is established: diabetic retinopathy can be asymptomatic, timely detection matters, and effective laser and injection treatments already exist. Vision outcomes take larger samples and longer follow-up than diagnostic studies. External care is difficult to observe. A same-day answer can remove delays even when the algorithm is not the only active ingredient.

The ACCESS trial and RAIDERS trial support that operational case. Immediate, point-of-care pathways can improve recorded follow-through in some settings. A product can be useful before every downstream question is answered.

But the burden changes with the claim. “Returns a diagnostic result at the point of care” can be supported by an accuracy and technical-yield study. “Improves referral” requires a credible comparator, aligned care stages and comparable follow-up capture and timing. “Prevents blindness” requires comparative longitudinal evidence that the program reduced vision loss or blindness, with treatment and follow-up denominators.

Programs large enough to report hundreds of thousands or millions of screens are also large enough to make the missing denominators consequential. If follow-up occurs outside the screening system, that is a reason to build linkage or sampling—not a reason to call the outcome known.

What proof would look like

The next retinal-AI study does not need another isolated sensitivity headline. It needs a linked ledger of unique people.

For every person with an image attempted, report:

  1. whether repeat capture or dilation was needed;
  2. whether the system returned a definitive result;
  3. whether the result was positive or ungradable, and the rule for each;
  4. whether referral was recommended;
  5. whether a confirmatory examination occurred, including outside the originating system;
  6. whether referable or vision-threatening disease was confirmed;
  7. whether treatment was indicated;
  8. whether treatment began;
  9. whether the intended treatment episode was completed; and
  10. visual acuity or another longitudinal vision outcome at a stated time.

Every transition needs a numerator, its immediate denominator and a time window. Repeated cameras and repeated visits need cohort identifiers so one person remains one person. Unknown follow-up should stay unknown. Workflow components—same-day results, counselling, scheduling, reminders, navigation, transport and specialist overread—should be recorded rather than credited silently to the classifier.

This is not an impossible standard. It is the ordinary accounting required to move from a test to a health outcome.

The bottom line

Can the AI eye exam read the retina? In many settings, yes. The diagnostic evidence is substantial, and our reconstructed AEYE matrices reproduce the reported sensitivity and specificity.

Can an AI-enabled pathway help more people reach eye care? Sometimes. The six-study pooled result is positive, but it is highly heterogeneous, denominator-sensitive and inseparable from the workflow around the model.

Can the current public evidence show that these programs complete treatment and preserve sight at scale? Not yet. Across 17 selected pathway rows representing 14,261 people, two of those rows independently report 10 treatment indications. Six report 56 treatment or management events, but only 33 have a named modality or a clearly documented first-treatment event. No selected row reports completion of the intended treatment course or a denominator for longitudinal vision outcome.

Zero measured vision outcomes is an evidence gap, not evidence of zero benefit.

The public record shows where the patient entered the funnel. It rarely shows where the patient emerged.

How this analysis was built

Ground Truth searched public evidence through 30 August 2026 and structured 29 unique primary or implementation studies into 41 arm or cohort rows across 14 countries. Seventeen studies entered the pathway-outcome census; six paired comparisons entered the referral synthesis. We separately mapped 15 current public claims to the furthest endpoint reported in their linked evidence. This was a citation-seeded evidence census, not a PRISMA-complete systematic review.

The protocol was frozen before the full source census but after several seed findings were known. It is not preregistered. We preserved unique-person lineage, patient-versus-eye units, image-attempt and definitive-result denominators, follow-up ascertainment, internal-versus-external capture, and treatment indication, initiation and completion as separate fields.

The review describes a Mantel–Haenszel analysis in R’s meta package. In that implementation, the random-effects summary uses inverse-variance weighting; with REML estimation and Hartung–Knapp confidence intervals, we reproduced the published risk ratio, risk difference and heterogeneity to the reported precision. We then ran stage-alignment, randomized-only, comparator-rule, risk-of-bias and leave-one-out analyses.

Three separately scoped research passes were reconciled against primary sources. The executable package validates 58 source records, 41 study rows, 15 claims, 36 aggregate facts and three diagnostic 2×2 tables. Thirteen gold tests pass. All 58 locally archived source files pass byte, hash and format checks. The AEYE pivotal Figure 2 is archived as its own hashed primary-source asset; the STATUS patient-flow figure is preserved inside a hashed supplementary ZIP, and the referral review’s DOCX and PDF supplements are separately archived and hashed.

The reusable data are published under CC BY 4.0: download the 17-study patient-count ledger, comparative referral effects, sensitivity analyses, diagnostic reconstructions, claim-endpoint map and source manifest.

This analysis is based entirely on the public record; no company or study-author outreach is part of the publication workflow. Corrections can be sent to corrections@groundtruth.health and will be logged publicly when they change the record.

Key sources

Disclosures & provenance

Published
31 Aug 2026
Author
The Ground Truth editor. Editorial standard →
Funding
Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
Data
Download the dataset · released under CC BY 4.0
Corrections
None to date. Corrections log → · Challenge this analysis