Analysis · Screening
The AI eye exam can read the retina. Can it save sight?
Retinal AI can identify referable diabetic retinopathy quickly and accurately. What happens next is much less certain. Ground Truth traced 29 studies and 15 public claims through the care chain. In the 17 selected patient-count rows, referral gains depended on the surrounding workflow; only two separately counted who needed treatment, and none reported completed treatment or later vision outcomes.
Evidence at a glance
29 studies · evidence by stage
- Diagnostic accuracy
- Strong, with conditions. Regulator-approved systems generally identify referable disease accurately; technical yield varies with camera, operator, repeat capture and dilation.
- Recorded referral follow-through
- Positive signal, pathway-dependent. The six-study pooled estimate favored AI-enabled pathways, but results varied sharply and every comparison bundled workflow changes with the algorithm.
- Completed treatment and preserved sight
- Evidence gap. Treatment completion and longitudinal vision remain unmeasured across all 17 selected patient-count pathways. The public record therefore leaves the frequency of benefit unknown.
Why stage-specific judgments? Ground Truth’s 1–5 ratings assess one attributable claim as stated. This article synthesizes multiple claims and outcomes, so evidence strength is reported by stage.
A participant repeated across two study counts still represents one person.
On 25 August, AEYE Health and academic collaborators published the results of three pivotal studies of an autonomous diabetic-retinopathy screening system. The paper’s title says the studies included “over 1,200 patients.” Its three displayed cohort counts add to 1,213: 531 in AEYE-1, 317 in AEYE-2 and 365 in AEYE-3. The 1,210 total can be reconstructed only by mixing cohort stages: 531 screened in AEYE-1, 317 enrolled in AEYE-2 and 362 enrolled in AEYE-3. The arithmetic is recoverable; the denominator mixes cohort stages.
The methods say all 317 people in AEYE-2 were a subset of the 531 in AEYE-1, tested again with a different camera. Add only the two independent cohorts, excluding the repeated-camera substudy, and the maximum number of unique people screened is 896. Using the displayed cohort sum as the denominator, at least 317 of the 1,213 camera-study entries, 26.1%, represent participants counted a second time.
The diagnostic result survives. The denominator needs a cohort-lineage qualifier.
The paper was funded by AEYE Health. Its disclosures list six authors as company employees and one additional author as a board member; the paper says the funder had no role in study design, data collection, analysis, interpretation or the decision to publish. The results remain intact; the funding and author relationships increase the importance of reproducible denominators and clear cohort accounting.
We reconstructed all three diagnostic matrices and reproduced the reported sensitivity and specificity. The cohort chart below shows the exact counts and reconstructed imageability fractions. Ground Truth derived those fractions from the Figure 2 analysis sets. They describe those analysis sets; cohort-flow and all-screened technical-yield denominators remain separate. Only these integers reproduce the paper’s estimates and Wilson intervals within those sets.
The system classified retinal images well. The second study adds camera-specific evidence by retesting 317 AEYE-1 participants. The abstract states that “Imageability was >99% in all studies,” and a separate company page repeats “>99% imageability.” Yet Section 3.3 prints AEYE-3 at exactly 99%, with a 95% interval running down to 96.97%. The only integer fraction within its Figure 2 analysis set that reproduces that interval is 331/335, or 98.81%, which is consistent with the paper’s rounded 99%. The exact underlying fractions are absent from the paper.
The accuracy result stands. This denominator problem is the first example of a pattern that runs through the public record on retinal AI: a claim may be supported at one rung of the care chain and then travel farther than the denominator underneath it.
Ground Truth’s AI-versus-clinician scoreboard previously highlighted a sensitivity advantage in Thailand. In the national evaluation, on the same images, Google’s deep-learning system had 91.4% sensitivity for vision-threatening diabetic retinopathy versus 84.8% for regional retina-specialist overreaders (p=.024), while specificity was effectively identical: 95.4% versus 95.5% (p=.98). We marked the patient-outcome field “No.” This investigation begins there.
This investigation asks both whether AI can read the retina and whether the person behind the image reaches confirmation, treatment and preserved sight.
One person, two camera-study entries
The headline sum combines camera-study entries. The lineage exposes repeated participation while preserving the camera-specific accuracy result.
1,213 camera-study entries are at most 896 unique screened people
View chart data
| Stage | What happens | Note |
|---|---|---|
| Three reported camera-study cohorts | AEYE-1 531 + AEYE-2 317 + AEYE-3 365 = 1,213 entries | Figure 2 builds the 1,210 total with 362 enrolled in AEYE-3 alongside screened counts, switching cohort stages from AEYE-3's 365 screened. |
| (paired lineage) AEYE-1 · 531 participants | Topcon NW400; contains the 317-person AEYE-2 paired-camera subset | AEYE-2 reuses AEYE-1 participants and provides paired-camera evidence from the same people. |
| (separate cohort) AEYE-3 · 365 participants | Aurora portable camera; separately enrolled cohort | |
| At most 896 unique screened people | 531 + 365; at least 317 of 1,213 entries (26.1%) repeat a participant |
- Repeated participation leaves the camera-specific accuracy estimates intact while limiting the unique-person total and the independence of AEYE-2.
- Figure 2 resolves the three-entry arithmetic gap: 531 + 317 + 362 = 1,210. The total mixes screened and enrolled stages; the repeated-person count remains unresolved.
The reconstructed camera-specific matrices reproduce the published metrics
Rows are reference-center more-than-mild diabetic-retinopathy status; columns are AEYE-DS status. Counts are shown as true positive, false negative, true negative and false positive.
| Camera study | Screened | True positive / false negative / true negative / false positive | Sensitivity | Specificity | Reconstructed imageability |
|---|---|---|---|---|---|
| AEYE-1 · Topcon NW400 | 531 | 53 / 4 / 370 / 35 | 92.98% | 91.36% | 462/466 · 99.14% |
| AEYE-2 · Aurora · paired subset | 317 | 34 / 3 / 233 / 16 | 91.89% | 93.57% | 286/288 · 99.31% |
| AEYE-3 · Aurora · separate cohort | 365 | 37 / 3 / 258 / 33 | 92.50% | 88.66% | 331/335 · 98.81% |
- The paper prints exact imageability point estimates for AEYE-1 and AEYE-2, a rounded 99% for AEYE-3, and Wilson intervals. The fractions themselves are absent. Ground Truth uniquely reconstructed them within the Figure 2 bounds.
- These fractions describe the Figure 2 analysis sets. AEYE-2 reuses 317 AEYE-1 participants and provides paired-camera evidence from the same people.
- The archived Figure 2 separately supports the three retained 2×2 matrices.
The part AI often does well
Diabetic retinopathy is a strong use case for medical AI. It is common, retinal photographs are standardized, and sight-threatening disease can be treated if people are found and cared for in time. An autonomous system can move image interpretation into a primary-care clinic, pharmacy or community program and return an answer while the patient is still there.
Across 82 studies covering 887,244 examinations and 25 regulator-approved systems, a 2025 systematic review reported pooled patient-level sensitivity of 93% and specificity of 90%. Those are high aggregate accuracy estimates. The public files omit the patient-level 2×2 cells needed for us to reproduce the pooled results independently, and performance—especially specificity—varies substantially across settings. The review nevertheless supports the narrower conclusion that regulator-approved retinal AI can identify referable disease from retinal photographs.
The condition is that the system first has to produce an answer.
“Imageability” reflects the full acquisition process: camera, operator, capture attempts, dilation policy and management of ungradable images. In a Mayo Clinic deployment, 580 of 1,052 people had AI-gradable photographs before dilation. After the protocol offered reflex dilation and repeat imaging for initially ungradable photographs, 965 ultimately had an AI-gradable set. In a Johns Hopkins deployment without reflex dilation, only 118 of 241 received a diagnostic output. A 2026 German evaluation reported 555 definitive results from 875 attempts under its no-retake, no-dilation workflow.
Taken together, these studies show that acquisition policy shapes performance. Controlled product comparisons would be needed to rank intrinsic imageability. A headline percentage that begins after unusable images have been excluded answers a different question from the proportion of all people who walked in and left with a result.
Diagnostic performance is the strongest link in the evidence chain. Technical yield and downstream care remain setting- and workflow-dependent.
The path from referral to treatment
The words around screening make the care chain sound short. A camera finds disease; the patient is referred; treatment prevents blindness.
The measurable chain is longer:
image attempted → definitive result → positive result → referral recommended → referral attended → disease confirmed → treatment indicated → treatment or management documented → treatment completed → vision measured
We built that ladder into a structured dataset and searched prospective and real-world diabetic-retinopathy AI studies published from 2020 through 30 August 2026. The pathway census contains 17 studies. For comparative studies, the numeric ledger selects the AI arm; otherwise it selects the single implementation cohort. That produces 17 pathway rows with a selected baseline of 14,261 people. The baseline is usually screening attempts when reported, then arm or cohort enrollment. One source-defined exception is explicit: Google Thailand uses its 7,651-person analysis cohort after 7,940 people were screened for inclusion.
We extracted every patient count each study reported. Because studies stop reporting at different stages, the rows draw on changing cohorts. Read each line independently as the minimum documented count among studies reporting that stage; sequential conversion requires a shared cohort.
Across changing contributor sets, the ledger records 3,049 definitive AI results, 918 cases with no definitive result or an ungradable image, 3,773 people eligible for referral, 630 recorded attendances, 477 confirmatory examinations and 90 confirmed disease cases. Eligible and attempted totals can exceed 14,261 because studies use different enrollment, screening and analysis denominators. The matching imageable and definitive-result totals come from the same eight studies, although the two concepts remain distinct. The pathway ledger below preserves the full reconciliation and every study row.
How many patients needed treatment? Only 10, across two studies, were explicitly counted as needing it. How many received treatment? Thirty-three patients across five studies had a named treatment or a clearly documented first-treatment event. A broader count reaches 56 across six studies by adding 23 Maine patients described only as receiving unspecified “appropriate treatments.” The 33 are included in the 56; the 10 come from a different study set and overlap only partly. These overlapping evidence layers come from different study sets and must remain separate.
Across the full 14,261-person census, those counts are minimum observed shares of 0.07%, 0.23% and 0.39%. Within the studies contributing each row, they equal 10/85, 33/289 and 56/334 of recorded attendees. Each percentage describes a different study set; sequential conversion would require a shared denominator. Many attendees appropriately required no treatment; only EyeArt directly observed indication followed by treatment, in three of six patients.
The Maine group also appears in the disease-confirmed count. Table 1 reports 23 of 45 follow-up patients as diabetic-retinopathy positive; the text says 51% “were found to have [diabetic retinopathy] and received appropriate treatments.” We therefore carry one reconstructed 23-person group in both evidence layers; a separately observed 23→23 transition is unavailable. The source leaves modality, timing, completion and the possible inclusion of observation undefined.
The missing data are concentrated. Google Thailand supplies 7,651 of the 14,261 selected patients, 53.6%. Rows with no attendance count supply 8,822 patients, 61.9% of the census; rows with no treatment or management count supply 10,267, 72.0%. These figures characterize only the selected public evidence. Generalization to an average program or representative patient population would be unwarranted.
Some apparent transitions use non-nested cohorts:
- Australia’s 74 AI referral flags and 28 clinical referrals use different criteria.
- EyeArt’s 92 referrals include all 53 inconclusive outputs alongside definitive results, breaking the subset relationship with its 127 definitive AI results.
- Excluding those rows leaves a valid definitive-result-to-referral link of 462/2,219 across five studies, 20.8%.
Other rows enter the ledger at different points:
- Google’s national Thailand paper reports 7,940 people screened for inclusion and a 7,651-person analysis cohort. The selected baseline is that analysis cohort. Its 2,412 composite referrals make up 63.9% of the 3,773 referral-eligible column sum and cover diabetic retinopathy, diabetic macular edema, ungradable images and poor visual acuity. The figure therefore represents a broader referral-eligibility category.
- RAIDERS contributes a 136-person post-positive randomized arm. DeepDR-LLM begins with 144 already referral-selected patients.
- Jordan contributes 402 screened people to the census, but its downstream percentage-derived counts conflict. They remain marked with question marks and excluded from stage totals and links.
We used strict definitions. Advice to seek care counts as a referral recommendation; attendance requires a recorded visit; confirmed disease requires a clinical diagnosis; and treatment delivery requires a documented treatment event. A first laser or injection establishes initiation. Completion requires evidence that the intended course ended. We excluded same-day specialist examinations among patients already inside an eye hospital from community referral uptake.
No study in the 17-study census documents completion of the indicated treatment plan or reports a denominator for longitudinal visual outcome. Treatment completion and sight preservation may have occurred; the public record leaves their frequency unknown.
The 17-study patient-count ledger
Every cell is an extracted canonical-stage patient count from one selected AI or single-pathway row. Both percentage columns are documented-minimum ratios; valid links use nested counts within rows.
What 17 selected pathway rows actually counted
14,261 patients across 17 selected pathway rows. Use an explicit source-defined analysis cohort where declared; otherwise use screening attempts when reported, then selected arm or cohort enrollment. Both percentage columns are minimum documented shares; incidence and complete ascertainment require linked denominators.
Treatment evidence, in three overlapping layers
| Evidence layer | People | Studies | Recorded minimum in all 14,261 | Count ÷ recorded attendees in the same studies | What the studies documented |
|---|---|---|---|---|---|
| Indication explicitly reported | 10 | 2/17 | 0.07% | 10/85 · 11.8% | Two studies separately counted who was judged to need treatment. |
| Specified treatment event | 33 | 5/17 | 0.23% | 33/289 · 11.4% | Named modality or clearly documented first-treatment event; contained within the inclusive 56. |
| Inclusive treatment / management event | 56 | 6/17 | 0.39% | 56/334 · 16.8% | Adds Maine’s same 23 people described only as receiving unspecified ‘appropriate treatments.’ |
Counts and percentages
Column sums combine different contributor cohorts. The contributor set and reporting-row baseline change by stage.
| Stage | Documented n | n / full 14,261 | Studies reporting | Reporting-row baselines | Documented n / reporting-row baselines |
|---|---|---|---|---|---|
| Eligible / enrolled | 14,646 | Not comparable | 16/17 | 13,941 | Not comparable |
| Screening attempted | 14,270 | Not comparable | 15/17 | 13,981 | Not comparable |
| Definitive AI result | 3,049 | 21.4% | 8/17 | 3,967 | 76.9% |
| No definitive result or ungradable | 918 | 6.44% | 8/17 | 3,967 | 23.1% |
| Study-defined AI referral flag | 977 | 6.85% | 10/17 | 4,995 | 19.6% |
| Study-defined referral eligible | 3,773 | 26.5% | 15/17 | 13,365 | 28.2% |
| Follow-up status ascertained | 1,094 | 7.67% | 11/17 | 4,919 | 22.2% |
| Recorded referral attendance | 630 | 4.42% | 13/17 | 5,439 | 11.6% |
| Confirmatory examination | 477 | 3.34% | 9/17 | 4,326 | 11.0% |
| Disease confirmed† | 90 | 0.63% | 3/17 | 1,208 | 7.45% |
| Treatment indication independently reported | 10 | 0.07% | 2/17 | 436 | 2.29% |
| Treatment or management event documented | 56 | 0.39% | 6/17 | 3,994 | 1.40% |
| Treatment course completed | — | Unknown | 0/17 | — | Unknown |
| Longitudinal vision outcome measured | — | Unknown | 0/17 | — | Unknown |
† Maine reports the same reconstructed 23-person group at both stages; the transition between them remains unobserved.
Valid within-row links
Explore every extracted patient-count row (17 studies)
Every extracted canonical-stage patient count
Symbol key: — unreported · ? conflicting percentage-only reports without an exact n · † source count shares a reported group across stages or lacks a separately enumerated indication; see its row note.
Screening and triage
| Study pathway row | Eligible | Attempted | Imageable | Definitive result | No definitive result | Referral flag | Referral eligible |
|---|---|---|---|---|---|---|---|
| STATUS · AIUnited States · baseline 2,243 (screening attempted)2,243 attempts → 1,459 definitive AI results + 784 without a definitive result or ungradable → 279 positive → 99 internal visits → 9 first-visit treatments. | 2,243 | 2,243 | 1,459 | 1,459 | 784 | 279 | 279 |
| Thailand platform · AIThailand · baseline 708 (screening attempted)201 AI flags → 129 over-reader-retained referrals → 115 attendances → 48 with vision-threatening diabetic retinopathy → 18 treated. | 708 | 708 | — | — | — | 201 | 129 |
| MogaIndia · baseline 343 (screening attempted) | 343 | 343 | — | — | — | — | 64 |
| Rural MaineUnited States · baseline 320 (screening attempted)The same 23-person diabetic-retinopathy-positive group is described as receiving unspecified ‘appropriate treatments.’ Ground Truth reconstructed 23 by linking the paper's 51%-of-45 sentence to Table 1's 23 diabetic-retinopathy-positive patients; modality, timing and whether this included observation are unreported. | — | 320 | — | — | — | 61 | 61 |
| Boothgarh · home AIIndia · baseline 200 (screening attempted)The 86 referrals included 36 diabetic-retinopathy-positive and 50 ungradable screens in the supplement; the linked gradability result is eye-level. | 200 | 200 | — | — | — | — | 86 |
| EyeArt follow-up · AIUnited States · baseline 180 (screening attempted)The 92 referrals include all 53 inconclusive outputs alongside definitive results, breaking the subset relationship with the 127 definitive AI results. | 185 | 180 | 127 | 127 | 53 | 39 | 92 |
| AEYE real-worldIsrael · baseline 256 (screening attempted) | 256 | 256 | 245 | 245 | 11 | 76 | 76 |
| Northern OntarioCanada · baseline 202 (screening attempted) | 202 | 202 | 189 | 189 | 13 | 42 | 42 |
| RAIDERS · AIRwanda · baseline 136 (eligible)Nested AI arm after 827 screened → 823 analyzed → 275 positive/randomized → 136 assigned AI. | 136 | — | — | — | — | — | 136 |
| ACCESS · AIUnited States · baseline 81 (screening attempted)Study lineage: 170 candidates → 164 randomized → 81 assigned to the selected AI row. | 81 | 81 | 81 | 81 | 0 | 25 | 25 |
| Northern IndiaIndia · baseline 390 (screening attempted) | 390 | 390 | — | — | — | — | 159 |
| AustraliaAustralia · baseline 236 (screening attempted)The 74 AI-positive outputs and 28 clinical referrals use different definitions; direct linkage is unavailable. | 456 | 236 | 232 | 232 | 4 | 74 | 28 |
| DeepDR-LLM · AIChina · baseline 144 (eligible)The 144-person baseline begins with an already referral-selected RDR cohort. | 144 | — | — | — | — | — | 144 |
| Google ThailandThailand · baseline 7,651 (analysis cohort)The source reports 7,940 screened for inclusion and 7,651 eligible for analysis. The selected census baseline is the 7,651-person analysis cohort. Its 2,412 composite referrals combine diabetic retinopathy, diabetic macular edema, ungradable images or poor visual acuity and represent a broader referral-eligibility category. | 7,651 | 7,940 | — | — | — | — | 2,412 |
| BelizeBelize · baseline 275 (screening attempted) | 639 | 275 | 245 | 245 | 30 | 40 | 40 |
| B-PRODUCTIVE · AIBangladesh · baseline 494 (screening attempted)Parent trial: 993 eligible → 494 AI / 499 control. Side branch: 140 positive + 23 insufficient → 163 same-day specialist exams at eye clinics. | 494 | 494 | 471 | 471 | 23 | 140 | — |
| Jordan pharmacyJordan · baseline 402 (screening attempted)Percentage-derived counts are 68 positive and 12 ungradable; reported referral completion implies 51/63, but the published chain conflicts and is quarantined. | 518 | 402 | — | — | — | 68? | ? |
| Column sums combine different contributor cohorts | 14,646 | 14,270 | 3,049 | 3,049 | 918 | 977 | 3,773 |
| Studies reporting | 16/17 | 15/17 | 8/17 | 8/17 | 8/17 | 10/17 | 15/17 |
Downstream care
| Study pathway row | Follow-up known | Attended | Confirm exam | Confirmed | Indication reported | Treatment / management | Completed | Vision outcome |
|---|---|---|---|---|---|---|---|---|
| STATUS · AIUnited States · baseline 2,243 (screening attempted) | 279 | 99 | 99 | — | — | 9 | — | — |
| Thailand platform · AIThailand · baseline 708 (screening attempted) | 129 | 115 | 115 | 48 | — | 18 | — | — |
| MogaIndia · baseline 343 (screening attempted) | 28 | 9 | — | — | — | 1 | — | — |
| Rural MaineUnited States · baseline 320 (screening attempted)† The same reconstructed 23 people supply both confirmed disease and unspecified treatment/management; modality, timing and active-treatment-versus-observation status remain unresolved. | — | 45 | 45 | 23† | — | 23† | — | — |
| Boothgarh · home AIIndia · baseline 200 (screening attempted)† Two people had recorded diabetic-retinopathy-treatment status; a separate treatment-indication count is unavailable. | — | 15 | 15 | — | — | 2† | — | — |
| EyeArt follow-up · AIUnited States · baseline 180 (screening attempted) | 92 | 51 | 51 | 19 | 6 | 3 | — | — |
| AEYE real-worldIsrael · baseline 256 (screening attempted) | 76 | 34 | 34 | — | 4 | — | — | — |
| Northern OntarioCanada · baseline 202 (screening attempted) | 42 | 32 | 32 | — | — | — | — | — |
| RAIDERS · AIRwanda · baseline 136 (eligible) | 136 | 70 | 70 | — | — | — | — | — |
| ACCESS · AIUnited States · baseline 81 (screening attempted) | 25 | 16 | 16 | — | — | — | — | — |
| Northern IndiaIndia · baseline 390 (screening attempted) | 128 | 23 | — | — | — | — | — | — |
| AustraliaAustralia · baseline 236 (screening attempted) | 15 | 9 | — | — | — | — | — | — |
| DeepDR-LLM · AIChina · baseline 144 (eligible) | 144 | 112 | — | — | — | — | — | — |
| Google ThailandThailand · baseline 7,651 (analysis cohort) | — | — | — | — | — | — | — | — |
| BelizeBelize · baseline 275 (screening attempted) | — | — | — | — | — | — | — | — |
| B-PRODUCTIVE · AIBangladesh · baseline 494 (screening attempted) | — | — | — | — | — | — | — | — |
| Jordan pharmacyJordan · baseline 402 (screening attempted) | ? | ? | — | — | — | — | — | — |
| Column sums combine different contributor cohorts | 1,094 | 630 | 477 | 90 | 10 | 56 | — | — |
| Studies reporting | 11/17 | 13/17 | 9/17 | 3/17 | 2/17 | 6/17 | 0/17 | 0/17 |
- The display selects one pathway row per study across a heterogeneous population. RAIDERS contributes a 136-person post-positive AI arm and DeepDR-LLM an already referral-selected 144-person cohort.
- Both percentage columns are minimum documented shares. Incidence and care-completion estimates require complete linked denominators. An omitted stage remains unknown.
- Stage totals use different contributor sets and remain separate. The 10, 33 and 56 treatment counts are overlapping evidence layers drawn from different study sets.
- Each link bar uses only rows with valid nested counts at both ends. EyeArt and Australia sit outside the definitive-result-to-referral link because their referral counts break the subset relationship with definitive AI results.
- Jordan contributes 402 screened people to the baseline, but its downstream percentage-derived counts conflict; they are shown with question marks and excluded from canonical stage totals and links.
- No study reported completed treatment or a denominator for longitudinal vision outcome. Those percentages remain unknown.
Sometimes the missing denominator nearly doubles the rate
Partial follow-up yields multiple defensible referral-attendance rates.
In an Australian implementation study, 28 people were referred, 15 could be contacted three months later, and nine said they had attended. That is 32.1% of everyone referred or 60% of the people reached. The second figure describes attendance among people whose status is known; the first is the conservative recorded minimum. The 13 unreachable people leave overall attendance uncertain.
The same distinction matters in Moga, India. Of 64 people referred, 28 were contacted one month later and nine said they had followed the advice and visited an ophthalmologist: 14.1% of everyone referred, or 32.1% of those contacted. We downgrade that to a reported referral attempt because five of the nine went to facilities without eye care.
The paper then lists 10 dispositions for nine people—one optical-coherence-tomography referral, those five visits, one eye-drop treatment, one recommendation for an anti–vascular endothelial growth factor eye injection, one laser treatment and one follow-up visit. The categories overlap or the report contains a counting error. We retain nine as the source-reported referral count; confirmed examinations remain uncounted. The laser qualifies as delivered diabetic-retinopathy treatment. Because the eye-drop indication is unspecified, we exclude it from diabetic-retinopathy treatment.
In the US STATUS program, 99 of 279 AI-positive patients had a directly observed Stanford ophthalmology visit within 90 days. The paper’s 69.2% any-provider estimate adds those internal visits, 35.5%, to 33.7% attributed to community visits. Ground Truth translates 33.7% of 279 to approximately 94 people; the paper omits both 94 and the exact extrapolation.
Outside follow-up was estimated from a phone sample: 152 people were called, 93 were reached and 68 of 93 reported follow-up. The paper labels 93/152 as 63.4%; the arithmetic is 61.2%. Our ledger retains the directly observed 99.
STATUS also had a nested AI–human salvage branch. Of 784 AI-ungradable cases, 664 entered human review; 561 were human-negative, 20 human-positive and 83 remained ungradable. Twelve of the paper’s own referrable denominator of 103 had an internal visit. The figure omits the other 120 AI-ungradable cases. The evidence record therefore shows the hybrid branch as contextual, nested data and keeps it outside the independent-cohort totals.
Internal records capture one care setting. For patients absent from them, attendance remains unknown because outside care may have occurred.
This sounds like bookkeeping because it is bookkeeping. It is also the difference between a care gap and a data gap.
The pooled effect survives a more complex interpretation
A 2026 systematic review and meta-analysis identified six comparative studies of AI-assisted diabetic-retinopathy pathways. We reconstructed all six 2×2 tables and reproduced the published result to the reported precision.
Across the six, recorded referral attendance was higher in the AI-enabled pathway: pooled risk ratio 1.89, with a 95% confidence interval from 1.18 to 3.03. The pooled absolute difference was 23.9 percentage points, from 12.8 to 35.1.
This is the strongest comparative evidence we found that AI-enabled pathways are associated with higher recorded referral attendance.
The estimate describes the observed referral-attendance difference for each bundled AI-enabled pathway, including its workflow changes; the algorithm-specific contribution remains unidentified.
The table below reproduces the review’s published-synthesis extraction:
| Comparative study | AI-pathway attendance | Comparator attendance | Risk ratio |
|---|---|---|---|
| ACCESS, United States | 16/25, 64.0% | 18/83, 21.7% | 2.95 |
| DeepDR-LLM, China | 112/144, 77.8% | 90/154, 58.4% | 1.33 |
| EyeArt follow-up, United States | 51/92, 55.4% | 182/974, 18.7% | 2.97 |
| RAIDERS, Rwanda | 70/136, 51.5% | 55/139, 39.6% | 1.30 |
| STATUS, United States | 99/279, 35.5% | 14/117, 12.0% | 2.97 |
| Thailand digital platform | 115/129, 89.1% | 124/175, 70.9% | 1.26 |
STATUS also reported 12 internal visits among 103 referral-eligible patients in its nested AI–human salvage branch. The table follows the review’s AI-versus-historical-control comparison; the evidence record includes the hybrid branch as contextual, nested data while keeping it outside the independent-cohort totals.
For ACCESS, the review uses 18/83 in the control arm. The primary analysis uses 18/82 after excluding one control participant enrolled in another screening study; our primary-endpoint sensitivity analysis uses 18/82.
ACCESS is also stage-asymmetric by design. The AI numerator is 16 of 25 participants with a disease-present result who completed a follow-up eye-care visit. The control numerator is 18 participants who completed an initial diabetic-eye examination after routine referral; all 18 received negative disease findings. That supports an end-to-end pathway contrast. A like-for-like referral-uptake comparison among screen-positive patients would require an equivalent disease-positive control denominator, which the study lacks.
The studies disagree sharply. I², a measure of between-study inconsistency, was 91.9%. The 95% prediction interval—our estimate of what a new similar setting might show—ran from a risk ratio of 0.53 to 6.78, spanning a possible decrease through a very large increase.
We also repeated the analysis while omitting one study at a time. With the review’s published rows, one of six confidence intervals crossed 1; after correcting the ACCESS control denominator and substituting Thailand’s closer stage comparison, three of six did.
The comparison arms differ too. In ACCESS, the control group received a routine referral recommendation while the AI group received a point-of-care result. In RAIDERS, the randomized contrast was an immediate AI result against delayed human grading. Other studies add counselling, scheduling, reminders or new referral rules. Our coding found at least two workflow differences in every comparison and as many as four; unreported components may raise that count.
Remove ACCESS and EyeArt—the two comparisons whose controls received routine or universal referral advice—and align Thailand to the same care stage in both arms. The remaining DeepDR-LLM, RAIDERS, STATUS and Thailand rows produce a descriptive pooled risk ratio of 1.47, with a confidence interval from 0.80 to 2.71. That restriction removes two of the three largest individual risk ratios. The two randomized trials alone give a risk ratio of 1.90 with a 95% confidence interval from 0.01 to 341.83, far too wide to guide a stable conclusion.
The pooled average favors AI-enabled pathways. Whether that result will travel to a new setting is unresolved.
At least two rows compare different clinical transitions across arms. ACCESS is one. Thailand is the clearest numerical mismatch.
In the implementation study, 115 of 129 AI-pathway true-positive referrals attended tertiary care. The review’s narrative compares that 89.1% with the primary paper’s stage-matched 17 of 22, or 77.3%, and reports p=.158. Yet its synthesis table uses 124 of 175, or 70.9%, for the manual period; that 124 appears to combine 116 intermediate confirmation visits and eight direct tertiary visits. The periods are observational and differ in care stage; the review nevertheless pooled the less aligned of its two manual-period numbers.
Correcting only that row barely changes the pooled point estimate: risk ratio 1.87, with a 95% confidence interval from 1.14 to 3.06. The result survives. The interpretation changes. The original rows estimate different clinical transitions, and the corrected prediction interval remains wide, 0.49 to 7.07.
The six AI denominators in the review’s main synthesis table sum to 805. Its supplementary summary-of-findings table reports “AI-assisted screening: 889,” 84 more. The manual-arm denominators reconcile exactly at 1,642. The main paper and its two supplements leave the extra 84 AI-side participants unexplained.
That internal discrepancy remains unresolved in the public report. Both totals describe study-defined groups entering different comparisons. A shared screening denominator remains unavailable.
On average, AI-enabled pathways recorded more referral attendance across whole workflows that bundled the algorithm with implementation changes. The review summarized the absolute difference as roughly one extra referral completion per four eligible people. Interpreting that difference as a screening-wide number needed to treat would be invalid because eligibility varied across studies, and our prediction interval for the absolute difference ran from −1.6 to +49.5 percentage points, crossing zero.
A positive mean, a wide range of plausible settings
Confidence and prediction intervals differ; workflows vary.
Recorded attendance rose on average, with wide variation
View chart data
| Comparison | Risk ratio | Limits | Interval type | Recorded attendance | Interval vs 1 | Design note |
|---|---|---|---|---|---|---|
| ACCESS | 2.95 | 1.78 to 4.88 | 95% confidence interval | 16/25 vs 18/83 | excludes 1 | Randomized controlled trial; low risk of bias; stage-asymmetric transition |
| DeepDR-LLM | 1.33 | 1.13 to 1.56 | 95% confidence interval | 112/144 vs 90/154 | excludes 1 | Sequential prospective comparison; moderate risk of bias |
| EyeArt follow-up | 2.97 | 2.37 to 3.72 | 95% confidence interval | 51/92 vs 182/974 | excludes 1 | Prospective cohort with historical comparator; serious risk of bias |
| RAIDERS | 1.30 | 1.00 to 1.69 | 95% confidence interval | 70/136 vs 55/139 | excludes 1 | Randomized controlled trial; low risk of bias |
| STATUS | 2.97 | 1.77 to 4.97 | 95% confidence interval | 99/279 vs 14/117 | excludes 1 | Historical workflow comparison; serious risk of bias |
| Thailand · published | 1.26 | 1.12 to 1.41 | 95% confidence interval | 115/129 vs 124/175 | excludes 1 | Alternating implementation periods; serious risk of bias |
| Thailand · same-stage | 1.15 | 0.91 to 1.46 | 95% confidence interval | 115/129 vs 17/22 | includes 1 | Alternative extraction from the same study; excluded from the six-comparison count |
| Published pooled | 1.89 | 1.18 to 3.03 | 95% confidence interval, adjusted for six studies | 6 studies | excludes 1 | Random effects; between-study inconsistency (I²) 91.9% |
| Same-stage pooled | 1.87 | 1.14 to 3.06 | 95% confidence interval, adjusted for six studies | Thailand alternate | excludes 1 | Sensitivity analysis; I² 90.9% |
| Published new-setting range | 1.89 | 0.53 to 6.78 | 95% prediction interval | I² 91.9% | includes 1 | 95% prediction interval for a new setting |
| Same-stage new-setting range | 1.87 | 0.49 to 7.07 | 95% prediction interval | Thailand alternate | includes 1 | Prediction interval after stage-aligning the Thailand row |
- Both the published and same-stage pooled results are low-certainty, highly heterogeneous and pathway-specific. Their 95% prediction intervals include lower attendance in a new setting.
- The same-stage Thailand point is an alternative extraction from the same study and stays outside the six-comparison count.
- ACCESS is stage-asymmetric by design: 16/25 is follow-up among AI disease-present patients, while 18/83 is an initial eye examination among the entire routinely referred control arm. A like-for-like uptake comparison among screen-positive patients is unavailable.
- The six rows use study-defined denominators and mixed designs. Their estimates apply to bundled pathways. The main-table AI denominators sum to 805; Supplementary Table 7 prints 889, an unexplained difference of 84.
- In the accessible table, confidence-interval limits are the lower and upper bounds around each estimate. Prediction intervals describe the wider range a new similar setting might show.
Every comparison changed more than the classifier
| Comparison | Instant result | Same-day option | Counselling | Scheduling | Reminders | Navigation / transport | Changed referral rule |
|---|---|---|---|---|---|---|---|
| ACCESS2 added; routine or universal referral | added in AI pathway | absent from AI-pathway documentation | documented in both pathways | absent from AI-pathway documentation | absent from AI-pathway documentation | absent from AI-pathway documentation | added in AI pathway |
| DeepDR-LLM3 added; manual grading workflow | added in AI pathway | absent from AI-pathway documentation | added in AI pathway | absent from AI-pathway documentation | absent from AI-pathway documentation | absent from AI-pathway documentation | added in AI pathway |
| EyeArt follow-up4 added; routine or universal referral | added in AI pathway | absent from AI-pathway documentation | added in AI pathway | added in AI pathway | absent from AI-pathway documentation | absent from AI-pathway documentation | added in AI pathway |
| RAIDERS2 added; manual grading workflow | added in AI pathway | added in AI pathway | documented in both pathways | absent from AI-pathway documentation | absent from AI-pathway documentation | documented in both pathways | absent from AI-pathway documentation |
| STATUS4 added; historical workflow | added in AI pathway | absent from AI-pathway documentation | added in AI pathway | added in AI pathway | absent from AI-pathway documentation | absent from AI-pathway documentation | added in AI pathway |
| Thailand3 added; manual grading workflow | added in AI pathway | absent from AI-pathway documentation | documented in both pathways | documented in both pathways | added in AI pathway | absent from AI-pathway documentation | added in AI pathway |
- A filled dot marks AI-pathway documentation absent from the comparator. Causal attribution for the attendance difference remains unresolved.
- Open dots identify shared counselling, scheduling or transport as co-interventions within both pathways.
Three funnels, three missing links
Three individual programs show why a valid patient journey must remain cohort-specific.
AEYE in an Israeli endocrinology clinic. In the 2026 real-world study, 256 people had an exam attempted, 245 received a definitive result and 76 were positive. The originating clinic documented 34 confirmatory examinations. Four of those 34 were judged to require treatment. For the other 42, outside care was unobserved, leaving confirmation status unknown. The paper leaves treatment initiation, completion and vision outcomes unreported.
Home AI screening in Boothgarh, India. In a three-arm pragmatic study, 200 people were enrolled in the home AI arm, 86 were referred and 15 reported by phone one month later that they had attended. Two participants had recorded diabetic-retinopathy treatment: one received laser, and one received laser plus an anti–vascular endothelial growth factor eye injection (supplementary Table 6). The intended treatment course, course completion and visual outcomes went unreported.
A linked publication from the same setting reported image gradability at eye level—224 of 362 analyzed eyes in the home/community AI arm. This different unit sits outside the patient-level flow. Conflicting arm-level phone ascertainment in the main article and supplement led us to use the full referral group as the attendance denominator.
A digital platform in Thailand. Here the chain extends farthest. Of 708 people screened in the AI period, 63 entered a separate visual-acuity referral branch. Among the remaining 645, AI flagged 201; specialist overread retained 129; 115 attended tertiary care; 48 had vision-threatening diabetic retinopathy; and 18 received laser or an eye injection.
The side branches matter. Of the 63 visual-acuity referrals, 52 attended. Overread rejected 72 AI positives, although 17 still presented, and found five false negatives, four of whom attended. We classify the two cataract treatments among 67 ungradable attenders separately from diabetic-retinopathy treatment. Completed treatment courses and vision outcomes went unreported.
The Thailand study documents high recorded attendance within a redesigned pathway. The program combined AI, specialist overread, reminders and active referral cancellation. The estimate therefore applies to the whole pathway; the algorithm-specific effect and superiority over the manual period remain unresolved.
Three programs, three separate missing links
The cascades use different gates, capture domains and endpoints. Each side branch remains separately visible.
AEYE real-world clinic: outside follow-up breaks the chain
View chart data
| Stage | What happens | Note |
|---|---|---|
| Exam attempted · 256 people | ||
| Definitive result · 245 | 95.7% of attempts | |
| Positive result · 76 | 31.0% of definitive results; 29.7% of attempts | |
| (observed internally) Internal confirmation · 34 | 44.7% of positive results | |
| (outside care unobserved) No internal confirmation record · 42 | Confirmation status unknown; outside care unobserved | |
| Treatment judged necessary · 4 | 11.8% of internally examined patients | |
| Initiation · completion · vision | Not reported |
- The 42 people without an internal confirmation record may include outside care. Their attendance status remains unknown.
- Treatment judged necessary records indication; initiation and completion remain unreported.
Boothgarh home AI: treatment status recorded for two; completion unknown
View chart data
| Stage | What happens | Note |
|---|---|---|
| Home AI arm · 200 people | ||
| Referred · 86 | 43.0% of the arm | |
| Recorded attendance at one month · 15 | 17.4% of referrals | The remaining 71 have no recorded visit; their attendance status remains unknown. |
| Participants with recorded diabetic-retinopathy treatment · 2 | One retinal laser; one retinal laser plus a drug injection into the eye (anti-VEGF) | Among 11 ungradable attenders, the supplement lists four as 'cat surgery,' three as 'cat Sx appointment' and four as no treatment. Completed cataract surgery remains unclear; these categories stay separate from diabetic-retinopathy treatment. |
| Completed diabetic-retinopathy episode · vision outcome | Not reported |
- Arm-level phone-ascertainment totals conflict between the main text and supplement, so attendance uses the full referral denominator.
- The linked 224/362 gradability count is eye-level and stays outside this patient-level funnel.
Thailand: the longest reported chain still has side branches
View chart data
| Stage | What happens | Note |
|---|---|---|
| Screened in AI period · 708 | ||
| (VA ≤20/70) Separate visual-acuity referral · 63 | 52 later presented at tertiary care | |
| (AI workflow denominator) Entered AI-image comparison · 645 | Denominator for the 201 AI positives: 645 | |
| AI referral-positive · 201 | 31.2% of 645 | Of 444 AI negatives, overread identified five false negatives; four attended after contact. |
| (true-positive referrals) Over-reader agreed · 129 | 64.2% of AI positives | |
| (false-positive branch) Over-reader disagreed · 72 | 17 still presented | 49 were reached to reverse referral (2 attended); 23 were not reached (15 attended). |
| True-positive referrals attending tertiary care · 115 | 89.1% of 129 | |
| Vision-threatening diabetic retinopathy confirmed · 48 | 45 diabetic macular edema; 3 severe retinopathy | |
| Retinal laser or eye injection received · 18 | 25 monitored; 5 referred onward | |
| Completed treatment episode · vision outcome | Not reported |
- The main spine follows over-reader-agreed true-positive referrals. The 52 VA-branch attenders, 17 false-positive presentations and four false-negative attenders remain separate side branches outside the 115 denominator.
- A separate 67-person ungradable branch included two cataract treatments, classified separately from diabetic-retinopathy treatment events.
- The pathway combined AI, specialist overread, reminders and referral reversal. The estimate applies to that full pathway; the algorithm-specific contribution remains unresolved.
A million screens, outcomes uncounted
The language of deployment moves more freely than the evidence.
In March, Google reported more than one million screenings through clinical partnerships across India, Thailand and Australia. The company says a diagnosis can arrive “in as little as two minutes” and “could potentially save their sight.” The sentence is appropriately hedged. The scale remains independently unverified, and the page leaves the unique-person question unanswered. The linked public studies lack one program-wide count moving from result to confirmation, completed treatment and vision.
Digital Diagnostics uses “Prevent blindness through early disease detection” as product language for LumineticsCore. The linked STATUS implementation study documents nine patients who “received treatment … at the first follow-up encounter.” That reaches treatment initiation, while treatment-plan completion and longitudinal vision remain unmeasured.
Orbis says patients given real-time AI diagnoses “are more likely to seek treatment.” The underlying RAIDERS randomized trial measured presentation for referral services within 30 days; treatment initiation fell outside its outcomes.
Other current scale statements have the same structural limit. Eyenuk has reported more than 230,000 patients screened.
Remidio uses the same 16-million figure for different stages: its product page displays “16M+ Patients Screened,” while an NITI Frontier Tech profile submitted under the Frontier Voices programme describes “16M+ people” risk-triaged and separately lists “1.2M+ retinal scans” analyzed. The profile says those scans generated epidemiological insights, including glaucoma-prevalence estimates.
India’s MadhuNetrAI announcement says 7,100 patients are “benefiting.” None of those public numbers, as currently linked, supplies a complete program-wide treatment and vision denominator.
Each statement may be accurate at its own stage. They describe distinct endpoints: screening volume, referral attendance, treatment initiation and preserved sight.
Each requires its own denominator.
The fair case for deploying before the final outcome
An accurate, regulated screening test can be deployed before a definitive population-level blindness trial is available.
The clinical logic is established: diabetic retinopathy can be asymptomatic, timely detection matters, and effective laser and injection treatments already exist. Vision outcomes take larger samples and longer follow-up than diagnostic studies. External care is difficult to observe. A same-day answer can remove delays as one component of a broader pathway.
The ACCESS trial and RAIDERS trial support that operational case. Immediate, point-of-care pathways can improve recorded follow-through in some settings. A product can be useful before every downstream question is answered.
The evidentiary burden rises with each claim. “Returns a diagnostic result at the point of care” can be supported by an accuracy and technical-yield study. “Improves referral” requires a credible comparator, aligned care stages and comparable follow-up capture and timing. “Prevents blindness” requires comparative longitudinal evidence that the program reduced vision loss or blindness, with treatment and follow-up denominators.
Programs large enough to report hundreds of thousands or millions of screens are also large enough to make the missing denominators consequential. Outside follow-up increases the need for linkage or representative sampling before claiming the outcome is known.
What proof would look like
The priority for the next retinal-AI study is a linked ledger of unique people.
For every person with an image attempted, report:
- whether repeat capture or dilation was needed;
- whether the system returned a definitive result;
- whether the result was positive or ungradable, and the rule for each;
- whether referral was recommended;
- whether a confirmatory examination occurred, including outside the originating system;
- whether referable or vision-threatening disease was confirmed;
- whether treatment was indicated;
- whether treatment began;
- whether the intended treatment episode was completed; and
- visual acuity or another longitudinal vision outcome at a stated time.
Every transition needs a numerator, its immediate denominator and a time window. Repeated cameras and repeated visits need cohort identifiers so one person remains one person. Unknown follow-up should stay unknown. Record workflow components—same-day results, counselling, scheduling, reminders, navigation, transport and specialist overread—explicitly so attribution reflects the whole pathway.
This standard is achievable: it is the ordinary accounting required to move from a test to a health outcome.
The bottom line
Can the AI eye exam read the retina? In many settings, yes. The diagnostic evidence is substantial, and our reconstructed AEYE matrices reproduce the reported sensitivity and specificity.
Can an AI-enabled pathway help more people reach eye care? Sometimes. The positive six-study pooled result is highly heterogeneous, denominator-sensitive and inseparable from the workflow around the model.
Can the current public evidence show that these programs complete treatment and preserve sight at scale? Current evidence stops short. Across 17 selected pathway rows representing 14,261 people, two of those rows independently report 10 treatment indications. Six rows report 56 treatment or management events. Five rows account for 33 patients with a named modality or clearly documented first-treatment event. No selected row reports completion of the intended treatment course or a denominator for longitudinal vision outcome.
The absence of measured vision outcomes leaves the effect on vision unknown.
The public record shows where the patient entered the funnel. It rarely shows where the patient emerged.
How this analysis was built
Ground Truth searched public evidence through 30 August 2026 and structured 29 unique primary or implementation studies into 41 arm or cohort rows across 14 countries. Seventeen studies entered the pathway-outcome census; six paired comparisons entered the referral synthesis. We separately mapped 15 current public claims to the furthest endpoint reported in their linked evidence. This analysis used a citation-seeded evidence-census protocol; PRISMA-complete systematic-review methods were outside its scope.
The protocol was frozen before the full source census but after several seed findings were known. The protocol lacks preregistration. We preserved unique-person lineage, patient-versus-eye units, image-attempt and definitive-result denominators, follow-up ascertainment, internal-versus-external capture, and treatment indication, initiation and completion as separate fields.
The review describes a Mantel–Haenszel analysis in R’s meta package. In that implementation, the random-effects summary uses inverse-variance weighting; with REML estimation and Hartung–Knapp confidence intervals, we reproduced the published risk ratio, risk difference and heterogeneity to the reported precision. We then ran stage-alignment, randomized-only, comparator-rule, risk-of-bias and leave-one-out analyses.
Three separately scoped research passes were reconciled against primary sources. The executable package validates 58 source records, 41 study rows, 15 claims, 36 aggregate facts and three diagnostic 2×2 tables. Thirteen gold tests pass. All 58 locally archived source files pass byte, hash and format checks. The AEYE pivotal Figure 2 is archived as its own hashed primary-source asset; the STATUS patient-flow figure is preserved inside a hashed supplementary ZIP, and the referral review’s DOCX and PDF supplements are separately archived and hashed.
The reusable data are published under CC BY 4.0: download the 17-study patient-count ledger, comparative referral effects, sensitivity analyses, diagnostic reconstructions, claim-endpoint map and source manifest.
This analysis is based entirely on the public record; no company or study-author outreach is part of the publication workflow. Corrections can be sent to corrections@groundtruth.health and will be logged publicly when they change the record.
Key sources
- AEYE pivotal synthesis: Frontiers in Digital Health, 25 August 2026
- AEYE real-world pathway: British Journal of Ophthalmology / PubMed
- Regulator-approved diagnostic-system review: npj Digital Medicine
- Referral-uptake systematic review and meta-analysis: npj Digital Medicine
- UK National Screening Committee external review: Automated grading in diabetic eye screening, 2026
- Thailand national diagnostic evaluation: The Lancet Digital Health / PubMed
- Thailand digital-platform implementation: Ophthalmology and Therapy / PMC
- ACCESS randomized trial: Nature Communications / PMC
- RAIDERS randomized trial: Ophthalmology Science / PMC
- STATUS implementation study: Clinical Ophthalmology / PMC
- EyeArt follow-up study: Ophthalmology Retina / PubMed
- Boothgarh pragmatic trial: Archives of Public Health / PMC
- Moga implementation study: JMIR Medical Informatics / PMC
- Australian implementation study: Scientific Reports / PMC
- Rural Maine EyeArt study: Journal of General Internal Medicine / PMC
Disclosures & provenance
- Published
- 31 Aug 2026 · last updated 1 Sep 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Data
- Download the dataset · released under CC BY 4.0
- Corrections
- None to date. Corrections log → · Challenge this analysis