Evidence at a glance

29 studies · evidence by stage

Diagnostic accuracy
Strong, with conditions. Regulator-approved systems generally identify referable disease accurately; technical yield varies with camera, operator, repeat capture and dilation.
Recorded referral follow-through
Positive signal, pathway-dependent. The six-study pooled estimate favored AI-enabled pathways, but results varied sharply and every comparison bundled workflow changes with the algorithm.
Completed treatment and preserved sight
Evidence gap. Treatment completion and longitudinal vision remain unmeasured across all 17 selected patient-count pathways. The public record therefore leaves the frequency of benefit unknown.

Why stage-specific judgments? Ground Truth’s 1–5 ratings assess one attributable claim as stated. This article synthesizes multiple claims and outcomes, so evidence strength is reported by stage.

A participant repeated across two study counts still represents one person.

On 25 August, AEYE Health and academic collaborators published the results of three pivotal studies of an autonomous diabetic-retinopathy screening system. The paper’s title says the studies included “over 1,200 patients.” Its three displayed cohort counts add to 1,213: 531 in AEYE-1, 317 in AEYE-2 and 365 in AEYE-3. The 1,210 total can be reconstructed only by mixing cohort stages: 531 screened in AEYE-1, 317 enrolled in AEYE-2 and 362 enrolled in AEYE-3. The arithmetic is recoverable; the denominator mixes cohort stages.

The methods say all 317 people in AEYE-2 were a subset of the 531 in AEYE-1, tested again with a different camera. Add only the two independent cohorts, excluding the repeated-camera substudy, and the maximum number of unique people screened is 896. Using the displayed cohort sum as the denominator, at least 317 of the 1,213 camera-study entries, 26.1%, represent participants counted a second time.

The diagnostic result survives. The denominator needs a cohort-lineage qualifier.

The paper was funded by AEYE Health. Its disclosures list six authors as company employees and one additional author as a board member; the paper says the funder had no role in study design, data collection, analysis, interpretation or the decision to publish. The results remain intact; the funding and author relationships increase the importance of reproducible denominators and clear cohort accounting.

We reconstructed all three diagnostic matrices and reproduced the reported sensitivity and specificity. The cohort chart below shows the exact counts and reconstructed imageability fractions. Ground Truth derived those fractions from the Figure 2 analysis sets. They describe those analysis sets; cohort-flow and all-screened technical-yield denominators remain separate. Only these integers reproduce the paper’s estimates and Wilson intervals within those sets.

The system classified retinal images well. The second study adds camera-specific evidence by retesting 317 AEYE-1 participants. The abstract states that “Imageability was >99% in all studies,” and a separate company page repeats “>99% imageability.” Yet Section 3.3 prints AEYE-3 at exactly 99%, with a 95% interval running down to 96.97%. The only integer fraction within its Figure 2 analysis set that reproduces that interval is 331/335, or 98.81%, which is consistent with the paper’s rounded 99%. The exact underlying fractions are absent from the paper.

The accuracy result stands. This denominator problem is the first example of a pattern that runs through the public record on retinal AI: a claim may be supported at one rung of the care chain and then travel farther than the denominator underneath it.

Ground Truth’s AI-versus-clinician scoreboard previously highlighted a sensitivity advantage in Thailand. In the national evaluation, on the same images, Google’s deep-learning system had 91.4% sensitivity for vision-threatening diabetic retinopathy versus 84.8% for regional retina-specialist overreaders (p=.024), while specificity was effectively identical: 95.4% versus 95.5% (p=.98). We marked the patient-outcome field “No.” This investigation begins there.

This investigation asks both whether AI can read the retina and whether the person behind the image reaches confirmation, treatment and preserved sight.

One person, two camera-study entries

The headline sum combines camera-study entries. The lineage exposes repeated participation while preserving the camera-specific accuracy result.

1,213 camera-study entries are at most 896 unique screened people

Three reported camera-study cohorts · AEYE-1 531 + AEYE-2 317 + AEYE-3 365 = 1,213 entriesThree reported camera-study cohortsAEYE-1 531 + AEYE-2 317 + AEYE-3 365 = 1,213entriesNoteFigure 2 builds the 1,210 total with 362enrolled in AEYE-3 alongside screened counts,switching cohort stages from AEYE-3's 365screened.AEYE-1 · 531 participants · Topcon NW400; contains the 317-person AEYE-2 paired-camera subsetAEYE-1 · 531participantsTopcon NW400;contains the317-person AEYE-2paired-camera subsetAEYE-3 · 365 participants · Aurora portable camera; separately enrolled cohortAEYE-3 · 365participantsAurora portablecamera; separatelyenrolled cohortNote — paired lineageAEYE-2 reuses AEYE-1 participants and providespaired-camera evidence from the same people.At most 896 unique screened people · 531 + 365; at least 317 of 1,213 entries (26.1%) repeat a participantAt most 896 unique screened people531 + 365; at least 317 of 1,213 entries (26.1%)repeat a participant
Source: AEYE pivotal synthesis, §3.1, Fig. 2 and §3.3/Table 3; Ground Truth cohort reconstruction.
View chart data
1,213 camera-study entries are at most 896 unique screened people flow data
StageWhat happensNote
Three reported camera-study cohortsAEYE-1 531 + AEYE-2 317 + AEYE-3 365 = 1,213 entriesFigure 2 builds the 1,210 total with 362 enrolled in AEYE-3 alongside screened counts, switching cohort stages from AEYE-3's 365 screened.
(paired lineage) AEYE-1 · 531 participantsTopcon NW400; contains the 317-person AEYE-2 paired-camera subsetAEYE-2 reuses AEYE-1 participants and provides paired-camera evidence from the same people.
(separate cohort) AEYE-3 · 365 participantsAurora portable camera; separately enrolled cohort
At most 896 unique screened people531 + 365; at least 317 of 1,213 entries (26.1%) repeat a participant
Read before citing
  • Repeated participation leaves the camera-specific accuracy estimates intact while limiting the unique-person total and the independence of AEYE-2.
  • Figure 2 resolves the three-entry arithmetic gap: 531 + 317 + 362 = 1,210. The total mixes screened and enrolled stages; the repeated-person count remains unresolved.

The reconstructed camera-specific matrices reproduce the published metrics

Rows are reference-center more-than-mild diabetic-retinopathy status; columns are AEYE-DS status. Counts are shown as true positive, false negative, true negative and false positive.

The reconstructed camera-specific matrices reproduce the published metrics table
Camera studyScreenedTrue positive / false negative / true negative / false positiveSensitivitySpecificityReconstructed imageability
AEYE-1 · Topcon NW40053153 / 4 / 370 / 3592.98%91.36%462/466 · 99.14%
AEYE-2 · Aurora · paired subset31734 / 3 / 233 / 1691.89%93.57%286/288 · 99.31%
AEYE-3 · Aurora · separate cohort36537 / 3 / 258 / 3392.50%88.66%331/335 · 98.81%
Source: AEYE pivotal synthesis, archived Figure 2, §3.3 and Table 3; Ground Truth diagnostic_2x2.csv and diagnostic_metrics.csv.
Read before citing
  • The paper prints exact imageability point estimates for AEYE-1 and AEYE-2, a rounded 99% for AEYE-3, and Wilson intervals. The fractions themselves are absent. Ground Truth uniquely reconstructed them within the Figure 2 bounds.
  • These fractions describe the Figure 2 analysis sets. AEYE-2 reuses 317 AEYE-1 participants and provides paired-camera evidence from the same people.
  • The archived Figure 2 separately supports the three retained 2×2 matrices.

The part AI often does well

Diabetic retinopathy is a strong use case for medical AI. It is common, retinal photographs are standardized, and sight-threatening disease can be treated if people are found and cared for in time. An autonomous system can move image interpretation into a primary-care clinic, pharmacy or community program and return an answer while the patient is still there.

Across 82 studies covering 887,244 examinations and 25 regulator-approved systems, a 2025 systematic review reported pooled patient-level sensitivity of 93% and specificity of 90%. Those are high aggregate accuracy estimates. The public files omit the patient-level 2×2 cells needed for us to reproduce the pooled results independently, and performance—especially specificity—varies substantially across settings. The review nevertheless supports the narrower conclusion that regulator-approved retinal AI can identify referable disease from retinal photographs.

The condition is that the system first has to produce an answer.

“Imageability” reflects the full acquisition process: camera, operator, capture attempts, dilation policy and management of ungradable images. In a Mayo Clinic deployment, 580 of 1,052 people had AI-gradable photographs before dilation. After the protocol offered reflex dilation and repeat imaging for initially ungradable photographs, 965 ultimately had an AI-gradable set. In a Johns Hopkins deployment without reflex dilation, only 118 of 241 received a diagnostic output. A 2026 German evaluation reported 555 definitive results from 875 attempts under its no-retake, no-dilation workflow.

Taken together, these studies show that acquisition policy shapes performance. Controlled product comparisons would be needed to rank intrinsic imageability. A headline percentage that begins after unusable images have been excluded answers a different question from the proportion of all people who walked in and left with a result.

Diagnostic performance is the strongest link in the evidence chain. Technical yield and downstream care remain setting- and workflow-dependent.

The path from referral to treatment

The words around screening make the care chain sound short. A camera finds disease; the patient is referred; treatment prevents blindness.

The measurable chain is longer:

image attempted → definitive result → positive result → referral recommended → referral attended → disease confirmed → treatment indicated → treatment or management documented → treatment completed → vision measured

We built that ladder into a structured dataset and searched prospective and real-world diabetic-retinopathy AI studies published from 2020 through 30 August 2026. The pathway census contains 17 studies. For comparative studies, the numeric ledger selects the AI arm; otherwise it selects the single implementation cohort. That produces 17 pathway rows with a selected baseline of 14,261 people. The baseline is usually screening attempts when reported, then arm or cohort enrollment. One source-defined exception is explicit: Google Thailand uses its 7,651-person analysis cohort after 7,940 people were screened for inclusion.

We extracted every patient count each study reported. Because studies stop reporting at different stages, the rows draw on changing cohorts. Read each line independently as the minimum documented count among studies reporting that stage; sequential conversion requires a shared cohort.

Across changing contributor sets, the ledger records 3,049 definitive AI results, 918 cases with no definitive result or an ungradable image, 3,773 people eligible for referral, 630 recorded attendances, 477 confirmatory examinations and 90 confirmed disease cases. Eligible and attempted totals can exceed 14,261 because studies use different enrollment, screening and analysis denominators. The matching imageable and definitive-result totals come from the same eight studies, although the two concepts remain distinct. The pathway ledger below preserves the full reconciliation and every study row.

How many patients needed treatment? Only 10, across two studies, were explicitly counted as needing it. How many received treatment? Thirty-three patients across five studies had a named treatment or a clearly documented first-treatment event. A broader count reaches 56 across six studies by adding 23 Maine patients described only as receiving unspecified “appropriate treatments.” The 33 are included in the 56; the 10 come from a different study set and overlap only partly. These overlapping evidence layers come from different study sets and must remain separate.

Across the full 14,261-person census, those counts are minimum observed shares of 0.07%, 0.23% and 0.39%. Within the studies contributing each row, they equal 10/85, 33/289 and 56/334 of recorded attendees. Each percentage describes a different study set; sequential conversion would require a shared denominator. Many attendees appropriately required no treatment; only EyeArt directly observed indication followed by treatment, in three of six patients.

The Maine group also appears in the disease-confirmed count. Table 1 reports 23 of 45 follow-up patients as diabetic-retinopathy positive; the text says 51% “were found to have [diabetic retinopathy] and received appropriate treatments.” We therefore carry one reconstructed 23-person group in both evidence layers; a separately observed 23→23 transition is unavailable. The source leaves modality, timing, completion and the possible inclusion of observation undefined.

The missing data are concentrated. Google Thailand supplies 7,651 of the 14,261 selected patients, 53.6%. Rows with no attendance count supply 8,822 patients, 61.9% of the census; rows with no treatment or management count supply 10,267, 72.0%. These figures characterize only the selected public evidence. Generalization to an average program or representative patient population would be unwarranted.

Some apparent transitions use non-nested cohorts:

  • Australia’s 74 AI referral flags and 28 clinical referrals use different criteria.
  • EyeArt’s 92 referrals include all 53 inconclusive outputs alongside definitive results, breaking the subset relationship with its 127 definitive AI results.
  • Excluding those rows leaves a valid definitive-result-to-referral link of 462/2,219 across five studies, 20.8%.

Other rows enter the ledger at different points:

  • Google’s national Thailand paper reports 7,940 people screened for inclusion and a 7,651-person analysis cohort. The selected baseline is that analysis cohort. Its 2,412 composite referrals make up 63.9% of the 3,773 referral-eligible column sum and cover diabetic retinopathy, diabetic macular edema, ungradable images and poor visual acuity. The figure therefore represents a broader referral-eligibility category.
  • RAIDERS contributes a 136-person post-positive randomized arm. DeepDR-LLM begins with 144 already referral-selected patients.
  • Jordan contributes 402 screened people to the census, but its downstream percentage-derived counts conflict. They remain marked with question marks and excluded from stage totals and links.

We used strict definitions. Advice to seek care counts as a referral recommendation; attendance requires a recorded visit; confirmed disease requires a clinical diagnosis; and treatment delivery requires a documented treatment event. A first laser or injection establishes initiation. Completion requires evidence that the intended course ended. We excluded same-day specialist examinations among patients already inside an eye hospital from community referral uptake.

No study in the 17-study census documents completion of the indicated treatment plan or reports a denominator for longitudinal visual outcome. Treatment completion and sight preservation may have occurred; the public record leaves their frequency unknown.

The 17-study patient-count ledger

Every cell is an extracted canonical-stage patient count from one selected AI or single-pathway row. Both percentage columns are documented-minimum ratios; valid links use nested counts within rows.

What 17 selected pathway rows actually counted

14,261 patients across 17 selected pathway rows. Use an explicit source-defined analysis cohort where declared; otherwise use screening attempts when reported, then selected arm or cohort enrollment. Both percentage columns are minimum documented shares; incidence and complete ascertainment require linked denominators.

Treatment evidence, in three overlapping layers

Three overlapping layers of reported treatment evidence
Evidence layerPeopleStudiesRecorded minimum in all 14,261Count ÷ recorded attendees in the same studiesWhat the studies documented
Indication explicitly reported102/170.07%10/85 · 11.8%Two studies separately counted who was judged to need treatment.
Specified treatment event335/170.23%33/289 · 11.4%Named modality or clearly documented first-treatment event; contained within the inclusive 56.
Inclusive treatment / management event566/170.39%56/334 · 16.8%Adds Maine’s same 23 people described only as receiving unspecified ‘appropriate treatments.’
Interpret these rows independently Only EyeArt independently observes indication followed by treatment: 3 of 6. The three rows above use different contributor sets and require independent interpretation.
Where the missing data sit Google Thailand supplies 7,651/14,261 patients (53.6%). Rows with no attendance count supply 8,822 (61.9%); rows with no treatment/management count supply 10,267 (72.0%). Missingness clusters in the larger rows.

Counts and percentages

Column sums combine different contributor cohorts. The contributor set and reporting-row baseline change by stage.

Stage counts and percentages across the 17-study pathway census
StageDocumented nn / full 14,261Studies reportingReporting-row baselinesDocumented n / reporting-row baselines
Eligible / enrolled14,646Not comparable16/1713,941Not comparable
Screening attempted14,270Not comparable15/1713,981Not comparable
Definitive AI result3,04921.4%8/173,96776.9%
No definitive result or ungradable9186.44%8/173,96723.1%
Study-defined AI referral flag9776.85%10/174,99519.6%
Study-defined referral eligible3,77326.5%15/1713,36528.2%
Follow-up status ascertained1,0947.67%11/174,91922.2%
Recorded referral attendance6304.42%13/175,43911.6%
Confirmatory examination4773.34%9/174,32611.0%
Disease confirmed†900.63%3/171,2087.45%
Treatment indication independently reported100.07%2/174362.29%
Treatment or management event documented560.39%6/173,9941.40%
Treatment course completedUnknown0/17Unknown
Longitudinal vision outcome measuredUnknown0/17Unknown

† Maine reports the same reconstructed 23-person group at both stages; the transition between them remains unobserved.

Valid within-row links

Explore every extracted patient-count row (17 studies)

Every extracted canonical-stage patient count

Symbol key: — unreported · ? conflicting percentage-only reports without an exact n · † source count shares a reported group across stages or lacks a separately enumerated indication; see its row note.

Screening and triage
Screening and triage patient counts in all 17 pathway studies
Study pathway rowEligibleAttemptedImageableDefinitive resultNo definitive resultReferral flagReferral eligible
STATUS · AIUnited States · baseline 2,243 (screening attempted)2,243 attempts → 1,459 definitive AI results + 784 without a definitive result or ungradable → 279 positive → 99 internal visits → 9 first-visit treatments.2,2432,2431,4591,459784279279
Thailand platform · AIThailand · baseline 708 (screening attempted)201 AI flags → 129 over-reader-retained referrals → 115 attendances → 48 with vision-threatening diabetic retinopathy → 18 treated.708708201129
MogaIndia · baseline 343 (screening attempted)34334364
Rural MaineUnited States · baseline 320 (screening attempted)The same 23-person diabetic-retinopathy-positive group is described as receiving unspecified ‘appropriate treatments.’ Ground Truth reconstructed 23 by linking the paper's 51%-of-45 sentence to Table 1's 23 diabetic-retinopathy-positive patients; modality, timing and whether this included observation are unreported.3206161
Boothgarh · home AIIndia · baseline 200 (screening attempted)The 86 referrals included 36 diabetic-retinopathy-positive and 50 ungradable screens in the supplement; the linked gradability result is eye-level.20020086
EyeArt follow-up · AIUnited States · baseline 180 (screening attempted)The 92 referrals include all 53 inconclusive outputs alongside definitive results, breaking the subset relationship with the 127 definitive AI results.185180127127533992
AEYE real-worldIsrael · baseline 256 (screening attempted)256256245245117676
Northern OntarioCanada · baseline 202 (screening attempted)202202189189134242
RAIDERS · AIRwanda · baseline 136 (eligible)Nested AI arm after 827 screened → 823 analyzed → 275 positive/randomized → 136 assigned AI.136136
ACCESS · AIUnited States · baseline 81 (screening attempted)Study lineage: 170 candidates → 164 randomized → 81 assigned to the selected AI row.8181818102525
Northern IndiaIndia · baseline 390 (screening attempted)390390159
AustraliaAustralia · baseline 236 (screening attempted)The 74 AI-positive outputs and 28 clinical referrals use different definitions; direct linkage is unavailable.45623623223247428
DeepDR-LLM · AIChina · baseline 144 (eligible)The 144-person baseline begins with an already referral-selected RDR cohort.144144
Google ThailandThailand · baseline 7,651 (analysis cohort)The source reports 7,940 screened for inclusion and 7,651 eligible for analysis. The selected census baseline is the 7,651-person analysis cohort. Its 2,412 composite referrals combine diabetic retinopathy, diabetic macular edema, ungradable images or poor visual acuity and represent a broader referral-eligibility category.7,6517,9402,412
BelizeBelize · baseline 275 (screening attempted)639275245245304040
B-PRODUCTIVE · AIBangladesh · baseline 494 (screening attempted)Parent trial: 993 eligible → 494 AI / 499 control. Side branch: 140 positive + 23 insufficient → 163 same-day specialist exams at eye clinics.49449447147123140
Jordan pharmacyJordan · baseline 402 (screening attempted)Percentage-derived counts are 68 positive and 12 ungradable; reported referral completion implies 51/63, but the published chain conflicts and is quarantined.51840268??
Column sums combine different contributor cohorts14,64614,2703,0493,0499189773,773
Studies reporting16/1715/178/178/178/1710/1715/17
Downstream care
Downstream care patient counts in all 17 pathway studies
Study pathway rowFollow-up knownAttendedConfirm examConfirmedIndication reportedTreatment / managementCompletedVision outcome
STATUS · AIUnited States · baseline 2,243 (screening attempted)27999999
Thailand platform · AIThailand · baseline 708 (screening attempted)1291151154818
MogaIndia · baseline 343 (screening attempted)2891
Rural MaineUnited States · baseline 320 (screening attempted)† The same reconstructed 23 people supply both confirmed disease and unspecified treatment/management; modality, timing and active-treatment-versus-observation status remain unresolved.454523†23†
Boothgarh · home AIIndia · baseline 200 (screening attempted)† Two people had recorded diabetic-retinopathy-treatment status; a separate treatment-indication count is unavailable.15152†
EyeArt follow-up · AIUnited States · baseline 180 (screening attempted)9251511963
AEYE real-worldIsrael · baseline 256 (screening attempted)7634344
Northern OntarioCanada · baseline 202 (screening attempted)423232
RAIDERS · AIRwanda · baseline 136 (eligible)1367070
ACCESS · AIUnited States · baseline 81 (screening attempted)251616
Northern IndiaIndia · baseline 390 (screening attempted)12823
AustraliaAustralia · baseline 236 (screening attempted)159
DeepDR-LLM · AIChina · baseline 144 (eligible)144112
Google ThailandThailand · baseline 7,651 (analysis cohort)
BelizeBelize · baseline 275 (screening attempted)
B-PRODUCTIVE · AIBangladesh · baseline 494 (screening attempted)
Jordan pharmacyJordan · baseline 402 (screening attempted)??
Column sums combine different contributor cohorts1,094630477901056
Studies reporting11/1713/179/173/172/176/170/170/17
Every displayed value is an extracted patient count. Column sums combine different contributor cohorts. Percentages use either the 14,261-patient selected-pathway census, reporting-row baselines, or valid nested counts for one labelled link.
Source: Ground Truth numeric pathway census; outputs/pathway_numeric_census.csv, pathway_numeric_stage_totals.csv and pathway_numeric_links.csv; public evidence searched through 30 Aug 2026.
Read before citing
  • The display selects one pathway row per study across a heterogeneous population. RAIDERS contributes a 136-person post-positive AI arm and DeepDR-LLM an already referral-selected 144-person cohort.
  • Both percentage columns are minimum documented shares. Incidence and care-completion estimates require complete linked denominators. An omitted stage remains unknown.
  • Stage totals use different contributor sets and remain separate. The 10, 33 and 56 treatment counts are overlapping evidence layers drawn from different study sets.
  • Each link bar uses only rows with valid nested counts at both ends. EyeArt and Australia sit outside the definitive-result-to-referral link because their referral counts break the subset relationship with definitive AI results.
  • Jordan contributes 402 screened people to the baseline, but its downstream percentage-derived counts conflict; they are shown with question marks and excluded from canonical stage totals and links.
  • No study reported completed treatment or a denominator for longitudinal vision outcome. Those percentages remain unknown.

Sometimes the missing denominator nearly doubles the rate

Partial follow-up yields multiple defensible referral-attendance rates.

In an Australian implementation study, 28 people were referred, 15 could be contacted three months later, and nine said they had attended. That is 32.1% of everyone referred or 60% of the people reached. The second figure describes attendance among people whose status is known; the first is the conservative recorded minimum. The 13 unreachable people leave overall attendance uncertain.

The same distinction matters in Moga, India. Of 64 people referred, 28 were contacted one month later and nine said they had followed the advice and visited an ophthalmologist: 14.1% of everyone referred, or 32.1% of those contacted. We downgrade that to a reported referral attempt because five of the nine went to facilities without eye care.

The paper then lists 10 dispositions for nine people—one optical-coherence-tomography referral, those five visits, one eye-drop treatment, one recommendation for an anti–vascular endothelial growth factor eye injection, one laser treatment and one follow-up visit. The categories overlap or the report contains a counting error. We retain nine as the source-reported referral count; confirmed examinations remain uncounted. The laser qualifies as delivered diabetic-retinopathy treatment. Because the eye-drop indication is unspecified, we exclude it from diabetic-retinopathy treatment.

In the US STATUS program, 99 of 279 AI-positive patients had a directly observed Stanford ophthalmology visit within 90 days. The paper’s 69.2% any-provider estimate adds those internal visits, 35.5%, to 33.7% attributed to community visits. Ground Truth translates 33.7% of 279 to approximately 94 people; the paper omits both 94 and the exact extrapolation.

Outside follow-up was estimated from a phone sample: 152 people were called, 93 were reached and 68 of 93 reported follow-up. The paper labels 93/152 as 63.4%; the arithmetic is 61.2%. Our ledger retains the directly observed 99.

STATUS also had a nested AI–human salvage branch. Of 784 AI-ungradable cases, 664 entered human review; 561 were human-negative, 20 human-positive and 83 remained ungradable. Twelve of the paper’s own referrable denominator of 103 had an internal visit. The figure omits the other 120 AI-ungradable cases. The evidence record therefore shows the hybrid branch as contextual, nested data and keeps it outside the independent-cohort totals.

Internal records capture one care setting. For patients absent from them, attendance remains unknown because outside care may have occurred.

This sounds like bookkeeping because it is bookkeeping. It is also the difference between a care gap and a data gap.

The pooled effect survives a more complex interpretation

A 2026 systematic review and meta-analysis identified six comparative studies of AI-assisted diabetic-retinopathy pathways. We reconstructed all six 2×2 tables and reproduced the published result to the reported precision.

Across the six, recorded referral attendance was higher in the AI-enabled pathway: pooled risk ratio 1.89, with a 95% confidence interval from 1.18 to 3.03. The pooled absolute difference was 23.9 percentage points, from 12.8 to 35.1.

This is the strongest comparative evidence we found that AI-enabled pathways are associated with higher recorded referral attendance.

The estimate describes the observed referral-attendance difference for each bundled AI-enabled pathway, including its workflow changes; the algorithm-specific contribution remains unidentified.

The table below reproduces the review’s published-synthesis extraction:

Referral attendance in six comparative studies
Comparative study AI-pathway attendance Comparator attendance Risk ratio
ACCESS, United States 16/25, 64.0% 18/83, 21.7% 2.95
DeepDR-LLM, China 112/144, 77.8% 90/154, 58.4% 1.33
EyeArt follow-up, United States 51/92, 55.4% 182/974, 18.7% 2.97
RAIDERS, Rwanda 70/136, 51.5% 55/139, 39.6% 1.30
STATUS, United States 99/279, 35.5% 14/117, 12.0% 2.97
Thailand digital platform 115/129, 89.1% 124/175, 70.9% 1.26

STATUS also reported 12 internal visits among 103 referral-eligible patients in its nested AI–human salvage branch. The table follows the review’s AI-versus-historical-control comparison; the evidence record includes the hybrid branch as contextual, nested data while keeping it outside the independent-cohort totals.

For ACCESS, the review uses 18/83 in the control arm. The primary analysis uses 18/82 after excluding one control participant enrolled in another screening study; our primary-endpoint sensitivity analysis uses 18/82.

ACCESS is also stage-asymmetric by design. The AI numerator is 16 of 25 participants with a disease-present result who completed a follow-up eye-care visit. The control numerator is 18 participants who completed an initial diabetic-eye examination after routine referral; all 18 received negative disease findings. That supports an end-to-end pathway contrast. A like-for-like referral-uptake comparison among screen-positive patients would require an equivalent disease-positive control denominator, which the study lacks.

The studies disagree sharply. I², a measure of between-study inconsistency, was 91.9%. The 95% prediction interval—our estimate of what a new similar setting might show—ran from a risk ratio of 0.53 to 6.78, spanning a possible decrease through a very large increase.

We also repeated the analysis while omitting one study at a time. With the review’s published rows, one of six confidence intervals crossed 1; after correcting the ACCESS control denominator and substituting Thailand’s closer stage comparison, three of six did.

The comparison arms differ too. In ACCESS, the control group received a routine referral recommendation while the AI group received a point-of-care result. In RAIDERS, the randomized contrast was an immediate AI result against delayed human grading. Other studies add counselling, scheduling, reminders or new referral rules. Our coding found at least two workflow differences in every comparison and as many as four; unreported components may raise that count.

Remove ACCESS and EyeArt—the two comparisons whose controls received routine or universal referral advice—and align Thailand to the same care stage in both arms. The remaining DeepDR-LLM, RAIDERS, STATUS and Thailand rows produce a descriptive pooled risk ratio of 1.47, with a confidence interval from 0.80 to 2.71. That restriction removes two of the three largest individual risk ratios. The two randomized trials alone give a risk ratio of 1.90 with a 95% confidence interval from 0.01 to 341.83, far too wide to guide a stable conclusion.

The pooled average favors AI-enabled pathways. Whether that result will travel to a new setting is unresolved.

At least two rows compare different clinical transitions across arms. ACCESS is one. Thailand is the clearest numerical mismatch.

In the implementation study, 115 of 129 AI-pathway true-positive referrals attended tertiary care. The review’s narrative compares that 89.1% with the primary paper’s stage-matched 17 of 22, or 77.3%, and reports p=.158. Yet its synthesis table uses 124 of 175, or 70.9%, for the manual period; that 124 appears to combine 116 intermediate confirmation visits and eight direct tertiary visits. The periods are observational and differ in care stage; the review nevertheless pooled the less aligned of its two manual-period numbers.

Correcting only that row barely changes the pooled point estimate: risk ratio 1.87, with a 95% confidence interval from 1.14 to 3.06. The result survives. The interpretation changes. The original rows estimate different clinical transitions, and the corrected prediction interval remains wide, 0.49 to 7.07.

The six AI denominators in the review’s main synthesis table sum to 805. Its supplementary summary-of-findings table reports “AI-assisted screening: 889,” 84 more. The manual-arm denominators reconcile exactly at 1,642. The main paper and its two supplements leave the extra 84 AI-side participants unexplained.

That internal discrepancy remains unresolved in the public report. Both totals describe study-defined groups entering different comparisons. A shared screening denominator remains unavailable.

On average, AI-enabled pathways recorded more referral attendance across whole workflows that bundled the algorithm with implementation changes. The review summarized the absolute difference as roughly one extra referral completion per four eligible people. Interpreting that difference as a screening-wide number needed to treat would be invalid because eligibility varied across studies, and our prediction interval for the absolute difference ran from −1.6 to +49.5 percentage points, crossing zero.

A positive mean, a wide range of plausible settings

Confidence and prediction intervals differ; workflows vary.

Recorded attendance rose on average, with wide variation

0.51248ACCESSACCESS: 2.95 (95% confidence interval 1.78 to 4.88; 16/25 vs 18/83)ACCESS: 2.95 (95% confidence interval 1.78 to 4.88; 16/25 vs 18/83)2.95 [1.78, 4.88]DeepDR-LLMDeepDR-LLM: 1.33 (95% confidence interval 1.13 to 1.56; 112/144 vs 90/154)DeepDR-LLM: 1.33 (95% confidence interval 1.13 to 1.56; 112/144 vs 90/154)1.33 [1.13, 1.56]EyeArt follow-upEyeArt follow-up: 2.97 (95% confidence interval 2.37 to 3.72; 51/92 vs 182/974)EyeArt follow-up: 2.97 (95% confidence interval 2.37 to 3.72; 51/92 vs 182/974)2.97 [2.37, 3.72]RAIDERSRAIDERS: 1.30 (95% confidence interval 1.00 to 1.69; 70/136 vs 55/139)RAIDERS: 1.30 (95% confidence interval 1.00 to 1.69; 70/136 vs 55/139)1.30 [1.00, 1.69]STATUSSTATUS: 2.97 (95% confidence interval 1.77 to 4.97; 99/279 vs 14/117)STATUS: 2.97 (95% confidence interval 1.77 to 4.97; 99/279 vs 14/117)2.97 [1.77, 4.97]Thailand · publishedThailand · published: 1.26 (95% confidence interval 1.12 to 1.41; 115/129 vs 124/175)Thailand · published: 1.26 (95% confidence interval 1.12 to 1.41; 115/129 vs 124/175)1.26 [1.12, 1.41]Thailand · same-stageThailand · same-stage: 1.15 (95% confidence interval 0.91 to 1.46; 115/129 vs 17/22) — interval includes 1Thailand · same-stage: 1.15 (95% confidence interval 0.91 to 1.46; 115/129 vs 17/22) — interval includes 11.15 [0.91, 1.46]Published pooledPublished pooled: 1.89 (95% confidence interval, adjusted for six studies 1.18 to 3.03; 6 studies)Published pooled: 1.89 (95% confidence interval, adjusted for six studies 1.18 to 3.03; 6 studies)1.89 [1.18, 3.03]Same-stage pooledSame-stage pooled: 1.87 (95% confidence interval, adjusted for six studies 1.14 to 3.06; Thailand alternate)Same-stage pooled: 1.87 (95% confidence interval, adjusted for six studies 1.14 to 3.06; Thailand alternate)1.87 [1.14, 3.06]Published new-setting rangePublished new-setting range: 1.89 (95% prediction interval 0.53 to 6.78; I² 91.9%) — interval includes 1Published new-setting range: 1.89 (95% prediction interval 0.53 to 6.78; I² 91.9%) — interval includes 11.89 [0.53, 6.78]Same-stage new-setting rangeSame-stage new-setting range: 1.87 (95% prediction interval 0.49 to 7.07; Thailand alternate) — interval includes 1Same-stage new-setting range: 1.87 (95% prediction interval 0.49 to 7.07; Thailand alternate) — interval includes 11.87 [0.49, 7.07]Risk ratio · comparator higher ← 1 → AI-enabled pathway higher
Circles show studies; diamonds, pooled estimates; amber, 95% prediction intervals.
Source: Ground Truth reconstruction of six published 2×2 referral tables; outputs/comparative_effects.csv and findings.json.
View chart data
Recorded attendance rose on average, with wide variation data table
ComparisonRisk ratioLimitsInterval typeRecorded attendanceInterval vs 1Design note
ACCESS2.951.78 to 4.8895% confidence interval16/25 vs 18/83excludes 1Randomized controlled trial; low risk of bias; stage-asymmetric transition
DeepDR-LLM1.331.13 to 1.5695% confidence interval112/144 vs 90/154excludes 1Sequential prospective comparison; moderate risk of bias
EyeArt follow-up2.972.37 to 3.7295% confidence interval51/92 vs 182/974excludes 1Prospective cohort with historical comparator; serious risk of bias
RAIDERS1.301.00 to 1.6995% confidence interval70/136 vs 55/139excludes 1Randomized controlled trial; low risk of bias
STATUS2.971.77 to 4.9795% confidence interval99/279 vs 14/117excludes 1Historical workflow comparison; serious risk of bias
Thailand · published1.261.12 to 1.4195% confidence interval115/129 vs 124/175excludes 1Alternating implementation periods; serious risk of bias
Thailand · same-stage1.150.91 to 1.4695% confidence interval115/129 vs 17/22includes 1Alternative extraction from the same study; excluded from the six-comparison count
Published pooled1.891.18 to 3.0395% confidence interval, adjusted for six studies6 studiesexcludes 1Random effects; between-study inconsistency (I²) 91.9%
Same-stage pooled1.871.14 to 3.0695% confidence interval, adjusted for six studiesThailand alternateexcludes 1Sensitivity analysis; I² 90.9%
Published new-setting range1.890.53 to 6.7895% prediction intervalI² 91.9%includes 195% prediction interval for a new setting
Same-stage new-setting range1.870.49 to 7.0795% prediction intervalThailand alternateincludes 1Prediction interval after stage-aligning the Thailand row
Read before citing
  • Both the published and same-stage pooled results are low-certainty, highly heterogeneous and pathway-specific. Their 95% prediction intervals include lower attendance in a new setting.
  • The same-stage Thailand point is an alternative extraction from the same study and stays outside the six-comparison count.
  • ACCESS is stage-asymmetric by design: 16/25 is follow-up among AI disease-present patients, while 18/83 is an initial eye examination among the entire routinely referred control arm. A like-for-like uptake comparison among screen-positive patients is unavailable.
  • The six rows use study-defined denominators and mixed designs. Their estimates apply to bundled pathways. The main-table AI denominators sum to 805; Supplementary Table 7 prints 889, an unexplained difference of 84.
  • In the accessible table, confidence-interval limits are the lower and upper bounds around each estimate. Prediction intervals describe the wider range a new similar setting might show.

Every comparison changed more than the classifier

Every comparison changed more than the classifier workflow matrix
ComparisonInstant resultSame-day optionCounsellingSchedulingRemindersNavigation / transportChanged referral rule
ACCESS2 added; routine or universal referraladded in AI pathwayabsent from AI-pathway documentationdocumented in both pathwaysabsent from AI-pathway documentationabsent from AI-pathway documentationabsent from AI-pathway documentationadded in AI pathway
DeepDR-LLM3 added; manual grading workflowadded in AI pathwayabsent from AI-pathway documentationadded in AI pathwayabsent from AI-pathway documentationabsent from AI-pathway documentationabsent from AI-pathway documentationadded in AI pathway
EyeArt follow-up4 added; routine or universal referraladded in AI pathwayabsent from AI-pathway documentationadded in AI pathwayadded in AI pathwayabsent from AI-pathway documentationabsent from AI-pathway documentationadded in AI pathway
RAIDERS2 added; manual grading workflowadded in AI pathwayadded in AI pathwaydocumented in both pathwaysabsent from AI-pathway documentationabsent from AI-pathway documentationdocumented in both pathwaysabsent from AI-pathway documentation
STATUS4 added; historical workflowadded in AI pathwayabsent from AI-pathway documentationadded in AI pathwayadded in AI pathwayabsent from AI-pathway documentationabsent from AI-pathway documentationadded in AI pathway
Thailand3 added; manual grading workflowadded in AI pathwayabsent from AI-pathway documentationdocumented in both pathwaysdocumented in both pathwaysadded in AI pathwayabsent from AI-pathway documentationadded in AI pathway
● added in AI pathway   ○ documented in both pathways   — absent from AI-pathway documentation
Source: Ground Truth extraction; outputs/workflow_components.csv and primary-study Methods/Results.
Read before citing
  • A filled dot marks AI-pathway documentation absent from the comparator. Causal attribution for the attendance difference remains unresolved.
  • Open dots identify shared counselling, scheduling or transport as co-interventions within both pathways.

Three individual programs show why a valid patient journey must remain cohort-specific.

AEYE in an Israeli endocrinology clinic. In the 2026 real-world study, 256 people had an exam attempted, 245 received a definitive result and 76 were positive. The originating clinic documented 34 confirmatory examinations. Four of those 34 were judged to require treatment. For the other 42, outside care was unobserved, leaving confirmation status unknown. The paper leaves treatment initiation, completion and vision outcomes unreported.

Home AI screening in Boothgarh, India. In a three-arm pragmatic study, 200 people were enrolled in the home AI arm, 86 were referred and 15 reported by phone one month later that they had attended. Two participants had recorded diabetic-retinopathy treatment: one received laser, and one received laser plus an anti–vascular endothelial growth factor eye injection (supplementary Table 6). The intended treatment course, course completion and visual outcomes went unreported.

A linked publication from the same setting reported image gradability at eye level—224 of 362 analyzed eyes in the home/community AI arm. This different unit sits outside the patient-level flow. Conflicting arm-level phone ascertainment in the main article and supplement led us to use the full referral group as the attendance denominator.

A digital platform in Thailand. Here the chain extends farthest. Of 708 people screened in the AI period, 63 entered a separate visual-acuity referral branch. Among the remaining 645, AI flagged 201; specialist overread retained 129; 115 attended tertiary care; 48 had vision-threatening diabetic retinopathy; and 18 received laser or an eye injection.

The side branches matter. Of the 63 visual-acuity referrals, 52 attended. Overread rejected 72 AI positives, although 17 still presented, and found five false negatives, four of whom attended. We classify the two cataract treatments among 67 ungradable attenders separately from diabetic-retinopathy treatment. Completed treatment courses and vision outcomes went unreported.

The Thailand study documents high recorded attendance within a redesigned pathway. The program combined AI, specialist overread, reminders and active referral cancellation. The estimate therefore applies to the whole pathway; the algorithm-specific effect and superiority over the manual period remain unresolved.

Three programs, three separate missing links

The cascades use different gates, capture domains and endpoints. Each side branch remains separately visible.

AEYE real-world clinic: outside follow-up breaks the chain

Exam attempted · 256 peopleExam attempted · 256 peopleDefinitive result · 245 · 95.7% of attemptsDefinitive result · 24595.7% of attemptsPositive result · 76 · 31.0% of definitive results; 29.7% of attemptsPositive result · 7631.0% of definitive results; 29.7% of attemptsInternal confirmation · 34 · 44.7% of positive resultsInternal confirmation· 3444.7% of positiveresultsNo internal confirmation record · 42 · Confirmation status unknown; outside care unobservedNo internalconfirmation record ·42Confirmation statusunknown; outside careunobservedTreatment judged necessary · 4 · 11.8% of internally examined patientsTreatment judged necessary · 411.8% of internally examined patientsInitiation · completion · vision · Not reportedInitiation · completion · visionNot reported
Source: AEYE 2026 Israeli endocrinology-clinic study, PMID 42373309, Methods and Results.
View chart data
AEYE real-world clinic: outside follow-up breaks the chain flow data
StageWhat happensNote
Exam attempted · 256 people
Definitive result · 24595.7% of attempts
Positive result · 7631.0% of definitive results; 29.7% of attempts
(observed internally) Internal confirmation · 3444.7% of positive results
(outside care unobserved) No internal confirmation record · 42Confirmation status unknown; outside care unobserved
Treatment judged necessary · 411.8% of internally examined patients
Initiation · completion · visionNot reported
Read before citing
  • The 42 people without an internal confirmation record may include outside care. Their attendance status remains unknown.
  • Treatment judged necessary records indication; initiation and completion remain unreported.

Boothgarh home AI: treatment status recorded for two; completion unknown

Home AI arm · 200 peopleHome AI arm · 200 peopleReferred · 86 · 43.0% of the armReferred · 8643.0% of the armRecorded attendance at one month · 15 · 17.4% of referralsRecorded attendance at one month · 1517.4% of referralsNoteThe remaining 71 have no recorded visit; theirattendance status remains unknown.Participants with recorded diabetic-retinopathy treatment · 2 · One retinal laser; one retinal laser plus a drug injection into the eye (anti-VEGF)Participants with recordeddiabetic-retinopathy treatment · 2One retinal laser; one retinal laser plus a druginjection into the eye (anti-VEGF)NoteAmong 11 ungradable attenders, the supplementlists four as 'cat surgery,' three as 'cat Sxappointment' and four as no treatment.Completed cataract surgery remains unclear;these categories stay separate fromdiabetic-retinopathy treatment.Completed diabetic-retinopathy episode · vision outcome · Not reportedCompleted diabetic-retinopathy episode ·vision outcomeNot reported
Source: Boothgarh pragmatic study, Results and supplement Tables 6–7; Ground Truth conservative recorded-minimum extraction.
View chart data
Boothgarh home AI: treatment status recorded for two; completion unknown flow data
StageWhat happensNote
Home AI arm · 200 people
Referred · 8643.0% of the arm
Recorded attendance at one month · 1517.4% of referralsThe remaining 71 have no recorded visit; their attendance status remains unknown.
Participants with recorded diabetic-retinopathy treatment · 2One retinal laser; one retinal laser plus a drug injection into the eye (anti-VEGF)Among 11 ungradable attenders, the supplement lists four as 'cat surgery,' three as 'cat Sx appointment' and four as no treatment. Completed cataract surgery remains unclear; these categories stay separate from diabetic-retinopathy treatment.
Completed diabetic-retinopathy episode · vision outcomeNot reported
Read before citing
  • Arm-level phone-ascertainment totals conflict between the main text and supplement, so attendance uses the full referral denominator.
  • The linked 224/362 gradability count is eye-level and stays outside this patient-level funnel.

Thailand: the longest reported chain still has side branches

Screened in AI period · 708Screened in AI period · 708Separate visual-acuity referral · 63 · 52 later presented at tertiary careSeparate visual-acuityreferral · 6352 later presented attertiary careEntered AI-image comparison · 645 · Denominator for the 201 AI positives: 645Entered AI-imagecomparison · 645Denominator for the201 AI positives: 645AI referral-positive · 201 · 31.2% of 645AI referral-positive · 20131.2% of 645NoteOf 444 AI negatives, overread identified fivefalse negatives; four attended after contact.Over-reader agreed · 129 · 64.2% of AI positivesOver-reader agreed ·12964.2% of AI positivesOver-reader disagreed · 72 · 17 still presentedOver-reader disagreed· 7217 still presentedNote — false-positive branch49 were reached to reverse referral (2attended); 23 were not reached (15 attended).True-positive referrals attending tertiary care · 115 · 89.1% of 129True-positive referrals attending tertiarycare · 11589.1% of 129Vision-threatening diabetic retinopathy confirmed · 48 · 45 diabetic macular edema; 3 severe retinopathyVision-threatening diabetic retinopathyconfirmed · 4845 diabetic macular edema; 3 severe retinopathyRetinal laser or eye injection received · 18 · 25 monitored; 5 referred onwardRetinal laser or eye injection received · 1825 monitored; 5 referred onwardCompleted treatment episode · vision outcome · Not reportedCompleted treatment episode · vision outcomeNot reported
Source: Thailand digital-platform study, Fig. 1, Tables 1–3 and Results; Ground Truth branch reconstruction.
View chart data
Thailand: the longest reported chain still has side branches flow data
StageWhat happensNote
Screened in AI period · 708
(VA ≤20/70) Separate visual-acuity referral · 6352 later presented at tertiary care
(AI workflow denominator) Entered AI-image comparison · 645Denominator for the 201 AI positives: 645
AI referral-positive · 20131.2% of 645Of 444 AI negatives, overread identified five false negatives; four attended after contact.
(true-positive referrals) Over-reader agreed · 12964.2% of AI positives
(false-positive branch) Over-reader disagreed · 7217 still presented49 were reached to reverse referral (2 attended); 23 were not reached (15 attended).
True-positive referrals attending tertiary care · 11589.1% of 129
Vision-threatening diabetic retinopathy confirmed · 4845 diabetic macular edema; 3 severe retinopathy
Retinal laser or eye injection received · 1825 monitored; 5 referred onward
Completed treatment episode · vision outcomeNot reported
Read before citing
  • The main spine follows over-reader-agreed true-positive referrals. The 52 VA-branch attenders, 17 false-positive presentations and four false-negative attenders remain separate side branches outside the 115 denominator.
  • A separate 67-person ungradable branch included two cataract treatments, classified separately from diabetic-retinopathy treatment events.
  • The pathway combined AI, specialist overread, reminders and referral reversal. The estimate applies to that full pathway; the algorithm-specific contribution remains unresolved.

A million screens, outcomes uncounted

The language of deployment moves more freely than the evidence.

In March, Google reported more than one million screenings through clinical partnerships across India, Thailand and Australia. The company says a diagnosis can arrive “in as little as two minutes” and “could potentially save their sight.” The sentence is appropriately hedged. The scale remains independently unverified, and the page leaves the unique-person question unanswered. The linked public studies lack one program-wide count moving from result to confirmation, completed treatment and vision.

Digital Diagnostics uses “Prevent blindness through early disease detection” as product language for LumineticsCore. The linked STATUS implementation study documents nine patients who “received treatment … at the first follow-up encounter.” That reaches treatment initiation, while treatment-plan completion and longitudinal vision remain unmeasured.

Orbis says patients given real-time AI diagnoses “are more likely to seek treatment.” The underlying RAIDERS randomized trial measured presentation for referral services within 30 days; treatment initiation fell outside its outcomes.

Other current scale statements have the same structural limit. Eyenuk has reported more than 230,000 patients screened.

Remidio uses the same 16-million figure for different stages: its product page displays “16M+ Patients Screened,” while an NITI Frontier Tech profile submitted under the Frontier Voices programme describes “16M+ people” risk-triaged and separately lists “1.2M+ retinal scans” analyzed. The profile says those scans generated epidemiological insights, including glaucoma-prevalence estimates.

India’s MadhuNetrAI announcement says 7,100 patients are “benefiting.” None of those public numbers, as currently linked, supplies a complete program-wide treatment and vision denominator.

Each statement may be accurate at its own stage. They describe distinct endpoints: screening volume, referral attendance, treatment initiation and preserved sight.

Each requires its own denominator.

The fair case for deploying before the final outcome

An accurate, regulated screening test can be deployed before a definitive population-level blindness trial is available.

The clinical logic is established: diabetic retinopathy can be asymptomatic, timely detection matters, and effective laser and injection treatments already exist. Vision outcomes take larger samples and longer follow-up than diagnostic studies. External care is difficult to observe. A same-day answer can remove delays as one component of a broader pathway.

The ACCESS trial and RAIDERS trial support that operational case. Immediate, point-of-care pathways can improve recorded follow-through in some settings. A product can be useful before every downstream question is answered.

The evidentiary burden rises with each claim. “Returns a diagnostic result at the point of care” can be supported by an accuracy and technical-yield study. “Improves referral” requires a credible comparator, aligned care stages and comparable follow-up capture and timing. “Prevents blindness” requires comparative longitudinal evidence that the program reduced vision loss or blindness, with treatment and follow-up denominators.

Programs large enough to report hundreds of thousands or millions of screens are also large enough to make the missing denominators consequential. Outside follow-up increases the need for linkage or representative sampling before claiming the outcome is known.

What proof would look like

The priority for the next retinal-AI study is a linked ledger of unique people.

For every person with an image attempted, report:

  1. whether repeat capture or dilation was needed;
  2. whether the system returned a definitive result;
  3. whether the result was positive or ungradable, and the rule for each;
  4. whether referral was recommended;
  5. whether a confirmatory examination occurred, including outside the originating system;
  6. whether referable or vision-threatening disease was confirmed;
  7. whether treatment was indicated;
  8. whether treatment began;
  9. whether the intended treatment episode was completed; and
  10. visual acuity or another longitudinal vision outcome at a stated time.

Every transition needs a numerator, its immediate denominator and a time window. Repeated cameras and repeated visits need cohort identifiers so one person remains one person. Unknown follow-up should stay unknown. Record workflow components—same-day results, counselling, scheduling, reminders, navigation, transport and specialist overread—explicitly so attribution reflects the whole pathway.

This standard is achievable: it is the ordinary accounting required to move from a test to a health outcome.

The bottom line

Can the AI eye exam read the retina? In many settings, yes. The diagnostic evidence is substantial, and our reconstructed AEYE matrices reproduce the reported sensitivity and specificity.

Can an AI-enabled pathway help more people reach eye care? Sometimes. The positive six-study pooled result is highly heterogeneous, denominator-sensitive and inseparable from the workflow around the model.

Can the current public evidence show that these programs complete treatment and preserve sight at scale? Current evidence stops short. Across 17 selected pathway rows representing 14,261 people, two of those rows independently report 10 treatment indications. Six rows report 56 treatment or management events. Five rows account for 33 patients with a named modality or clearly documented first-treatment event. No selected row reports completion of the intended treatment course or a denominator for longitudinal vision outcome.

The absence of measured vision outcomes leaves the effect on vision unknown.

The public record shows where the patient entered the funnel. It rarely shows where the patient emerged.

How this analysis was built

Ground Truth searched public evidence through 30 August 2026 and structured 29 unique primary or implementation studies into 41 arm or cohort rows across 14 countries. Seventeen studies entered the pathway-outcome census; six paired comparisons entered the referral synthesis. We separately mapped 15 current public claims to the furthest endpoint reported in their linked evidence. This analysis used a citation-seeded evidence-census protocol; PRISMA-complete systematic-review methods were outside its scope.

The protocol was frozen before the full source census but after several seed findings were known. The protocol lacks preregistration. We preserved unique-person lineage, patient-versus-eye units, image-attempt and definitive-result denominators, follow-up ascertainment, internal-versus-external capture, and treatment indication, initiation and completion as separate fields.

The review describes a Mantel–Haenszel analysis in R’s meta package. In that implementation, the random-effects summary uses inverse-variance weighting; with REML estimation and Hartung–Knapp confidence intervals, we reproduced the published risk ratio, risk difference and heterogeneity to the reported precision. We then ran stage-alignment, randomized-only, comparator-rule, risk-of-bias and leave-one-out analyses.

Three separately scoped research passes were reconciled against primary sources. The executable package validates 58 source records, 41 study rows, 15 claims, 36 aggregate facts and three diagnostic 2×2 tables. Thirteen gold tests pass. All 58 locally archived source files pass byte, hash and format checks. The AEYE pivotal Figure 2 is archived as its own hashed primary-source asset; the STATUS patient-flow figure is preserved inside a hashed supplementary ZIP, and the referral review’s DOCX and PDF supplements are separately archived and hashed.

The reusable data are published under CC BY 4.0: download the 17-study patient-count ledger, comparative referral effects, sensitivity analyses, diagnostic reconstructions, claim-endpoint map and source manifest.

This analysis is based entirely on the public record; no company or study-author outreach is part of the publication workflow. Corrections can be sent to corrections@groundtruth.health and will be logged publicly when they change the record.

Key sources

Disclosures & provenance

Published
31 Aug 2026 · last updated 1 Sep 2026
Author
The Ground Truth editor. Editorial standard →
Funding
Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
Data
Download the dataset · released under CC BY 4.0
Corrections
None to date. Corrections log → · Challenge this analysis