Analysis · Screening
The AI eye exam can read the retina. Can it save sight?
Retinal AI can identify referable diabetic retinopathy quickly and accurately. What happens next is much less certain. Ground Truth traced 29 studies and 15 public claims through the care chain. In the 17 selected patient-count rows, referral gains depended on the surrounding workflow; only two separately counted who needed treatment, and none reported completed treatment or later vision outcomes.
One person can appear in two study counts without becoming two people.
On 25 August, AEYE Health and academic collaborators published the results of three pivotal studies of an autonomous diabetic-retinopathy screening system. The paper’s title says the studies included “over 1,200 patients.” Its three displayed cohort counts add to 1,213: 531 in AEYE-1, 317 in AEYE-2 and 365 in AEYE-3. The 1,210 total can be reconstructed only by mixing cohort stages: 531 screened in AEYE-1, 317 enrolled in AEYE-2 and 362 enrolled in AEYE-3. The arithmetic is recoverable, but the denominator is not stage-consistent.
But the methods say all 317 people in AEYE-2 were a subset of the 531 in AEYE-1, tested again with a different camera. Add the two independent cohorts—not the repeated camera substudy—and the maximum number of unique people screened is 896. Using the displayed cohort sum as the denominator, at least 317 of the 1,213 camera-study entries, 26.1%, represent participants counted a second time.
The diagnostic result survives. The denominator needs a cohort-lineage qualifier.
The paper was funded by AEYE Health. Its disclosures list six authors as company employees and one additional author as a board member; the paper says the funder was not involved in study design, data collection, analysis, interpretation or the decision to publish. That does not negate the results, but it raises the importance of a reproducible denominator and clear cohort accounting.
We reconstructed all three diagnostic matrices and reproduced the reported sensitivity and specificity. The cohort chart below shows the exact counts and reconstructed imageability fractions. Those fractions are Ground Truth reconstructions—not printed cohort-flow counts or all-screened technical-yield fractions—and are the only integers within the Figure 2 analysis sets that reproduce the paper’s estimates and Wilson intervals.
The system classified retinal images well. The second study adds useful camera evidence, but not 317 new people or an independent replication. The abstract states that “Imageability was >99% in all studies,” and a separate company page repeats “>99% imageability.” Yet Section 3.3 prints AEYE-3 at exactly 99%, with a 95% interval running down to 96.97%. The only integer fraction within its Figure 2 analysis set that reproduces that interval is 331/335, or 98.81%, which is consistent with the paper’s rounded 99%. The paper does not supply the exact underlying fractions.
This is not an accuracy takedown. It is the first example of the problem that runs through the public record on retinal AI: a claim may be supported at one rung of the care chain and then travel farther than the denominator underneath it.
Ground Truth’s AI-versus-clinician scoreboard previously highlighted a sensitivity advantage in Thailand. In the national evaluation, on the same images, Google’s deep-learning system had 91.4% sensitivity for vision-threatening diabetic retinopathy versus 84.8% for regional retina-specialist overreaders (p=.024), while specificity was effectively identical: 95.4% versus 95.5% (p=.98). We marked the patient-outcome field “No.” This investigation begins there.
The question is no longer only whether AI can read the retina. It is whether the person behind the image reaches confirmation, treatment and preserved sight.
One person, two camera-study entries
The headline cohort sum is not a unique-person count. The lineage keeps repeated participation visible without erasing the camera-specific accuracy result.
1,213 camera-study entries are at most 896 unique screened people
View chart data
| Stage | What happens | Note |
|---|---|---|
| Three reported camera-study cohorts | AEYE-1 531 + AEYE-2 317 + AEYE-3 365 = 1,213 entries | Figure 2 shows that the 1,210 total uses 362 enrolled in AEYE-3 rather than its 365 screened, switching cohort stages. |
| (paired lineage) AEYE-1 · 531 participants | Topcon NW400; contains the 317-person AEYE-2 paired-camera subset | AEYE-2 reuses AEYE-1 participants; it is not an independent replication. |
| (separate cohort) AEYE-3 · 365 participants | Aurora portable camera; separately enrolled cohort | |
| At most 896 unique screened people | 531 + 365; at least 317 of 1,213 entries (26.1%) repeat a participant |
- Repeated participation does not invalidate the camera-specific accuracy estimates. It limits the unique-person total and the independence of AEYE-2.
- Figure 2 resolves the three-entry arithmetic gap: 531 + 317 + 362 = 1,210. The total mixes screened and enrolled stages and does not resolve the repeated-person count.
The camera-specific diagnostic matrices still reproduce
Rows are reference-center more-than-mild diabetic-retinopathy status; columns are AEYE-DS status. Counts are shown as true positive, false negative, true negative and false positive.
| Camera study | Screened | True positive / false negative / true negative / false positive | Sensitivity | Specificity | Reconstructed imageability |
|---|---|---|---|---|---|
| AEYE-1 · Topcon NW400 | 531 | 53 / 4 / 370 / 35 | 92.98% | 91.36% | 462/466 · 99.14% |
| AEYE-2 · Aurora · paired subset | 317 | 34 / 3 / 233 / 16 | 91.89% | 93.57% | 286/288 · 99.31% |
| AEYE-3 · Aurora · separate cohort | 365 | 37 / 3 / 258 / 33 | 92.50% | 88.66% | 331/335 · 98.81% |
- The paper prints exact imageability point estimates for AEYE-1 and AEYE-2, a rounded 99% for AEYE-3, and Wilson intervals, but no fractions. Ground Truth uniquely reconstructed the fractions within the Figure 2 bounds.
- The fractions are not all-screened technical-yield counts. AEYE-2 reuses 317 AEYE-1 participants and is camera evidence, not an independent-patient replication.
- The archived Figure 2 separately supports the three retained 2×2 matrices.
The part AI often does well
Diabetic retinopathy is a strong use case for medical AI. It is common, retinal photographs are standardized, and sight-threatening disease can be treated if people are found and cared for in time. An autonomous system can move image interpretation into a primary-care clinic, pharmacy or community program and return an answer while the patient is still there.
Across 82 studies covering 887,244 examinations and 25 regulator-approved systems, a 2025 systematic review reported pooled patient-level sensitivity of 93% and specificity of 90%. Those are high aggregate accuracy estimates. The public files do not include the patient-level 2×2 cells needed for us to reproduce the pooled results independently, and performance—especially specificity—varies substantially across settings. The review nevertheless supports the narrower conclusion that regulator-approved retinal AI can identify referable disease from retinal photographs.
The condition is that the system first has to produce an answer.
“Imageability” is not a property of software alone. It changes with the camera, the operator, the number of capture attempts, whether pupils are dilated, and what happens to an ungradable image. In a Mayo Clinic deployment, 580 of 1,052 people had AI-gradable photographs before dilation. After the protocol offered reflex dilation and repeat imaging for initially ungradable photographs, 965 ultimately had an AI-gradable set. In a Johns Hopkins deployment without reflex dilation, only 118 of 241 received a diagnostic output. A 2026 German evaluation reported 555 definitive results from 875 attempts under its no-retake, no-dilation workflow.
Those studies do not establish that one product is intrinsically more imageable than another. They show that acquisition policy is part of performance. A headline percentage that begins after unusable images have been excluded answers a different question from the proportion of all people who walked in and left with a result.
Diagnostic performance is the strongest link in the evidence chain. Technical yield and downstream care remain setting- and workflow-dependent.
A referral is not treatment
The words around screening make the care chain sound short. A camera finds disease; the patient is referred; treatment prevents blindness.
The measurable chain is longer:
image attempted → definitive result → positive result → referral recommended → referral attended → disease confirmed → treatment indicated → treatment or management documented → treatment completed → vision measured
We built that ladder into a structured dataset and searched prospective and real-world diabetic-retinopathy AI studies published from 2020 through 30 August 2026. The pathway census contains 17 studies. For comparative studies, the numeric ledger selects the AI arm; otherwise it selects the single implementation cohort. That produces 17 pathway rows with a selected baseline of 14,261 people. The baseline is usually screening attempts when reported, then arm or cohort enrollment. One source-defined exception is explicit: Google Thailand uses its 7,651-person analysis cohort after 7,940 people were screened for inclusion.
We extracted every patient count each study reported. Because studies stop reporting at different stages, no single group of 14,261 people can be followed from the first row to the last. Read each line independently: it is the minimum documented count among studies reporting that stage, not a conversion rate.
Across changing contributor sets, the ledger records 3,049 definitive AI results, 918 cases with no definitive result or an ungradable image, 3,773 people eligible for referral, 630 recorded attendances, 477 confirmatory examinations and 90 confirmed disease cases. Eligible and attempted totals can exceed 14,261 because studies use different enrollment, screening and analysis denominators. The matching imageable and definitive-result totals come from the same eight studies, although the two concepts remain distinct. The pathway ledger below preserves the full reconciliation and every study row.
How many patients needed treatment? Only 10, across two studies, were explicitly counted as needing it. How many received treatment? Thirty-three patients across five studies had a named treatment or a clearly documented first-treatment event. A broader count reaches 56 across six studies by adding 23 Maine patients described only as receiving unspecified “appropriate treatments.” The 33 are included in the 56; the 10 come from a different study set and overlap only partly. These are overlapping evidence layers, not steps in one funnel.
Across the full 14,261-person census, those counts are minimum observed shares of 0.07%, 0.23% and 0.39%. Within the studies contributing each row, they equal 10/85, 33/289 and 56/334 of recorded attendees. Because each row uses a different study set, those percentages are not sequential conversion rates. Many attendees appropriately did not need treatment; only EyeArt directly observed indication followed by treatment, in three of six patients.
The Maine group also appears in the disease-confirmed count. Table 1 reports 23 of 45 follow-up patients as diabetic-retinopathy positive; the text says 51% “were found to have [diabetic retinopathy] and received appropriate treatments.” We therefore carry one reconstructed 23-person group in both evidence layers—not a separately observed 23→23 transition. The source does not define modality, timing, completion or whether “appropriate treatments” included observation.
The missing data are concentrated. Google Thailand supplies 7,651 of the 14,261 selected patients, 53.6%. Rows with no attendance count supply 8,822 patients, 61.9% of the census; rows with no treatment or management count supply 10,267, 72.0%. These figures describe the selected public evidence—not an average program or representative patient population.
Some apparent transitions are not nested:
- Australia’s 74 AI referral flags and 28 clinical referrals use different criteria.
- EyeArt’s 92 referrals include all 53 inconclusive outputs, so they are not a subset of its 127 definitive AI results.
- Excluding those rows leaves a valid definitive-result-to-referral link of 462/2,219 across five studies, 20.8%.
Other rows enter the ledger at different points:
- Google’s national Thailand paper reports 7,940 people screened for inclusion and a 7,651-person analysis cohort. The selected baseline is that analysis cohort. Its 2,412 composite referrals make up 63.9% of the 3,773 referral-eligible column sum, but include diabetic retinopathy, diabetic macular edema, ungradable images and poor visual acuity. They are not an AI-positive count.
- RAIDERS contributes a 136-person post-positive randomized arm. DeepDR-LLM begins with 144 already referral-selected patients.
- Jordan contributes 402 screened people to the census, but its downstream percentage-derived counts conflict. They remain marked with question marks and excluded from stage totals and links.
We used strict definitions. Advice to seek care is not attendance. Attendance is not confirmed disease. “Treatment required” is not treatment delivered. A first laser or injection is evidence of initiation, not proof that the intended course was completed. Same-day specialist examinations among patients already inside an eye hospital were not recoded as community referral uptake.
No study in the 17-study census documents completion of the indicated treatment plan or reports a denominator for longitudinal visual outcome. That does not mean nobody completed treatment or preserved sight. It means the public studies we found cannot tell us how often either happened.
The 17-study patient-count ledger
Every cell is an extracted canonical-stage patient count from one selected AI or single-pathway row. Both percentage columns are documented-minimum ratios; valid links use nested counts within rows.
What 17 selected pathway rows actually counted
14,261 patients across 17 selected pathway rows. Use an explicit source-defined analysis cohort where declared; otherwise use screening attempts when reported, then selected arm or cohort enrollment. Both percentage columns are minimum documented shares, not incidence or complete-ascertainment estimates.
Treatment evidence, in three overlapping layers
| Evidence layer | People | Studies | Recorded minimum in all 14,261 | Count ÷ recorded attendees in the same studies | What the studies documented |
|---|---|---|---|---|---|
| Indication explicitly reported | 10 | 2/17 | 0.07% | 10/85 · 11.8% | Two studies separately counted who was judged to need treatment. |
| Specified treatment event | 33 | 5/17 | 0.23% | 33/289 · 11.4% | Named modality or clearly documented first-treatment event; contained within the inclusive 56. |
| Inclusive treatment / management event | 56 | 6/17 | 0.39% | 56/334 · 16.8% | Adds Maine’s same 23 people described only as receiving unspecified ‘appropriate treatments.’ |
Counts and percentages
Column sums—not a common cohort. The contributor set and reporting-row baseline change by stage.
| Stage | Documented n | n / full 14,261 | Studies reporting | Reporting-row baselines | Documented n / reporting-row baselines |
|---|---|---|---|---|---|
| Eligible / enrolled | 14,646 | Not comparable | 16/17 | 13,941 | Not comparable |
| Screening attempted | 14,270 | Not comparable | 15/17 | 13,981 | Not comparable |
| Definitive AI result | 3,049 | 21.4% | 8/17 | 3,967 | 76.9% |
| No definitive result or ungradable | 918 | 6.44% | 8/17 | 3,967 | 23.1% |
| Study-defined AI referral flag | 977 | 6.85% | 10/17 | 4,995 | 19.6% |
| Study-defined referral eligible | 3,773 | 26.5% | 15/17 | 13,365 | 28.2% |
| Follow-up status ascertained | 1,094 | 7.67% | 11/17 | 4,919 | 22.2% |
| Recorded referral attendance | 630 | 4.42% | 13/17 | 5,439 | 11.6% |
| Confirmatory examination | 477 | 3.34% | 9/17 | 4,326 | 11.0% |
| Disease confirmed† | 90 | 0.63% | 3/17 | 1,208 | 7.45% |
| Treatment indication independently reported | 10 | 0.07% | 2/17 | 436 | 2.29% |
| Treatment or management event documented | 56 | 0.39% | 6/17 | 3,994 | 1.40% |
| Treatment course completed | — | Unknown | 0/17 | — | Unknown |
| Longitudinal vision outcome measured | — | Unknown | 0/17 | — | Unknown |
† Maine supplies the same reconstructed 23-person group to disease confirmed and the inclusive treatment/management count; this is one reported group, not two consecutive events.
Valid within-row links
Explore every extracted patient-count row (17 studies)
Every extracted canonical-stage patient count
Symbol key: — not reported · ? conflicting percentage-only reports without an exact n · † source count shares a reported group across stages or lacks a separately enumerated indication; see its row note.
Screening and triage
| Study pathway row | Eligible | Attempted | Imageable | Definitive result | No definitive result | Referral flag | Referral eligible |
|---|---|---|---|---|---|---|---|
| STATUS · AIUnited States · baseline 2,243 (screening attempted)2,243 attempts → 1,459 definitive AI results + 784 without a definitive result or ungradable → 279 positive → 99 internal visits → 9 first-visit treatments. | 2,243 | 2,243 | 1,459 | 1,459 | 784 | 279 | 279 |
| Thailand platform · AIThailand · baseline 708 (screening attempted)201 AI flags → 129 over-reader-retained referrals → 115 attendances → 48 with vision-threatening diabetic retinopathy → 18 treated. | 708 | 708 | — | — | — | 201 | 129 |
| MogaIndia · baseline 343 (screening attempted) | 343 | 343 | — | — | — | — | 64 |
| Rural MaineUnited States · baseline 320 (screening attempted)The same 23-person diabetic-retinopathy-positive group is described as receiving unspecified ‘appropriate treatments.’ Ground Truth reconstructed 23 by linking the paper's 51%-of-45 sentence to Table 1's 23 diabetic-retinopathy-positive patients; modality, timing and whether this included observation are unreported. | — | 320 | — | — | — | 61 | 61 |
| Boothgarh · home AIIndia · baseline 200 (screening attempted)The 86 referrals included 36 diabetic-retinopathy-positive and 50 ungradable screens in the supplement; the linked gradability result is eye-level. | 200 | 200 | — | — | — | — | 86 |
| EyeArt follow-up · AIUnited States · baseline 180 (screening attempted)The 92 referrals include all 53 inconclusive outputs, so they are not a subset of the 127 definitive AI results. | 185 | 180 | 127 | 127 | 53 | 39 | 92 |
| AEYE real-worldIsrael · baseline 256 (screening attempted) | 256 | 256 | 245 | 245 | 11 | 76 | 76 |
| Northern OntarioCanada · baseline 202 (screening attempted) | 202 | 202 | 189 | 189 | 13 | 42 | 42 |
| RAIDERS · AIRwanda · baseline 136 (eligible)Nested AI arm after 827 screened → 823 analyzed → 275 positive/randomized → 136 assigned AI. | 136 | — | — | — | — | — | 136 |
| ACCESS · AIUnited States · baseline 81 (screening attempted)Study lineage: 170 candidates → 164 randomized → 81 assigned to the selected AI row. | 81 | 81 | 81 | 81 | 0 | 25 | 25 |
| Northern IndiaIndia · baseline 390 (screening attempted) | 390 | 390 | — | — | — | — | 159 |
| AustraliaAustralia · baseline 236 (screening attempted)The 74 AI-positive outputs and 28 clinical referrals use different definitions; they are not a direct link. | 456 | 236 | 232 | 232 | 4 | 74 | 28 |
| DeepDR-LLM · AIChina · baseline 144 (eligible)The 144-person baseline is an already referral-selected RDR cohort, not an all-screened population. | 144 | — | — | — | — | — | 144 |
| Google ThailandThailand · baseline 7,651 (analysis cohort)The source reports 7,940 screened for inclusion and 7,651 eligible for analysis. The selected census baseline is the 7,651-person analysis cohort. Its 2,412 composite referrals combine diabetic retinopathy, diabetic macular edema, ungradable images or poor visual acuity; they are not an AI-positive count. | 7,651 | 7,940 | — | — | — | — | 2,412 |
| BelizeBelize · baseline 275 (screening attempted) | 639 | 275 | 245 | 245 | 30 | 40 | 40 |
| B-PRODUCTIVE · AIBangladesh · baseline 494 (screening attempted)Parent trial: 993 eligible → 494 AI / 499 control. Side branch: 140 positive + 23 insufficient → 163 same-day specialist exams at eye clinics. | 494 | 494 | 471 | 471 | 23 | 140 | — |
| Jordan pharmacyJordan · baseline 402 (screening attempted)Percentage-derived counts are 68 positive and 12 ungradable; reported referral completion implies 51/63, but the published chain conflicts and is quarantined. | 518 | 402 | — | — | — | 68? | ? |
| Column sum—not a common cohort | 14,646 | 14,270 | 3,049 | 3,049 | 918 | 977 | 3,773 |
| Studies reporting | 16/17 | 15/17 | 8/17 | 8/17 | 8/17 | 10/17 | 15/17 |
Downstream care
| Study pathway row | Follow-up known | Attended | Confirm exam | Confirmed | Indication reported | Treatment / management | Completed | Vision outcome |
|---|---|---|---|---|---|---|---|---|
| STATUS · AIUnited States · baseline 2,243 (screening attempted) | 279 | 99 | 99 | — | — | 9 | — | — |
| Thailand platform · AIThailand · baseline 708 (screening attempted) | 129 | 115 | 115 | 48 | — | 18 | — | — |
| MogaIndia · baseline 343 (screening attempted) | 28 | 9 | — | — | — | 1 | — | — |
| Rural MaineUnited States · baseline 320 (screening attempted)† The same reconstructed 23 people supply both confirmed disease and unspecified treatment/management; the paper does not establish modality, timing, or that active treatment rather than observation was intended. | — | 45 | 45 | 23† | — | 23† | — | — |
| Boothgarh · home AIIndia · baseline 200 (screening attempted)† Two people had recorded diabetic-retinopathy-treatment status, but treatment indication was not separately enumerated. | — | 15 | 15 | — | — | 2† | — | — |
| EyeArt follow-up · AIUnited States · baseline 180 (screening attempted) | 92 | 51 | 51 | 19 | 6 | 3 | — | — |
| AEYE real-worldIsrael · baseline 256 (screening attempted) | 76 | 34 | 34 | — | 4 | — | — | — |
| Northern OntarioCanada · baseline 202 (screening attempted) | 42 | 32 | 32 | — | — | — | — | — |
| RAIDERS · AIRwanda · baseline 136 (eligible) | 136 | 70 | 70 | — | — | — | — | — |
| ACCESS · AIUnited States · baseline 81 (screening attempted) | 25 | 16 | 16 | — | — | — | — | — |
| Northern IndiaIndia · baseline 390 (screening attempted) | 128 | 23 | — | — | — | — | — | — |
| AustraliaAustralia · baseline 236 (screening attempted) | 15 | 9 | — | — | — | — | — | — |
| DeepDR-LLM · AIChina · baseline 144 (eligible) | 144 | 112 | — | — | — | — | — | — |
| Google ThailandThailand · baseline 7,651 (analysis cohort) | — | — | — | — | — | — | — | — |
| BelizeBelize · baseline 275 (screening attempted) | — | — | — | — | — | — | — | — |
| B-PRODUCTIVE · AIBangladesh · baseline 494 (screening attempted) | — | — | — | — | — | — | — | — |
| Jordan pharmacyJordan · baseline 402 (screening attempted) | ? | ? | — | — | — | — | — | — |
| Column sum—not a common cohort | 1,094 | 630 | 477 | 90 | 10 | 56 | — | — |
| Studies reporting | 11/17 | 13/17 | 9/17 | 3/17 | 2/17 | 6/17 | 0/17 | 0/17 |
- This is one selected pathway row per study, not a homogeneous screened population. RAIDERS contributes a 136-person post-positive AI arm and DeepDR-LLM an already referral-selected 144-person cohort.
- Both percentage columns are minimum documented shares, not incidence or care-completion estimates. An omitted stage remains unknown, not zero.
- Stage totals use different contributor sets and must not be connected as attrition. The 10, 33 and 56 treatment counts are overlapping evidence layers, not consecutive gates.
- Each link bar uses only rows with valid nested counts at both ends. EyeArt and Australia are excluded from the definitive-result-to-referral link because their referral counts are not subsets of definitive AI results.
- Jordan contributes 402 screened people to the baseline, but its downstream percentage-derived counts conflict; they are shown with question marks and excluded from canonical stage totals and links.
- No study reported completed treatment or a denominator for longitudinal vision outcome. Those percentages are unknown, not 0%.
Sometimes the missing denominator nearly doubles the rate
Even “attended referral” is not one number when researchers cannot observe every patient.
In an Australian implementation study, 28 people were referred, 15 could be contacted three months later, and nine said they had attended. That is 32.1% of everyone referred or 60% of the people reached. The second figure describes attendance among people whose status is known; the first is the conservative recorded minimum. Neither reveals what happened to the 13 people who could not be reached.
The same distinction matters in Moga, India. Of 64 people referred, 28 were contacted one month later and nine said they had followed the advice and visited an ophthalmologist: 14.1% of everyone referred, or 32.1% of those contacted. We downgrade that to a reported referral attempt because five of the nine went to facilities without eye care.
The paper then lists 10 dispositions for nine people—one optical-coherence-tomography referral, those five visits, one eye-drop treatment, one recommendation for an anti–vascular endothelial growth factor eye injection, one laser treatment and one follow-up visit. The categories overlap or the report contains a counting error. We retain nine as the source-reported referral count, not nine confirmed examinations. Only the laser is coded as delivered diabetic-retinopathy treatment; the eye drops are not identified as diabetic-retinopathy directed.
In the US STATUS program, 99 of 279 AI-positive patients had a directly observed Stanford ophthalmology visit within 90 days. The paper’s 69.2% any-provider estimate adds those internal visits, 35.5%, to 33.7% attributed to community visits. Ground Truth translates 33.7% of 279 to approximately 94 people, but the paper does not print 94 or show the exact extrapolation.
Outside follow-up was estimated from a phone sample: 152 people were called, 93 were reached and 68 of 93 reported follow-up. The paper labels 93/152 as 63.4%; the arithmetic is 61.2%. Our ledger retains the directly observed 99.
STATUS also had a nested AI–human salvage branch. Of 784 AI-ungradable cases, 664 entered human review; 561 were human-negative, 20 human-positive and 83 remained ungradable. Twelve of the paper’s own referrable denominator of 103 had an internal visit. The figure does not account for the other 120 AI-ungradable cases, so we show the hybrid branch as context instead of adding it as an independent cohort.
Internal records cannot rule out outside care. A patient absent from that record may have gone nowhere or may have gone elsewhere. “Not observed here” is not the same as “did not attend.”
This sounds like bookkeeping because it is bookkeeping. It is also the difference between a care gap and a data gap.
The pooled effect survives. Its simplicity does not
A 2026 systematic review and meta-analysis identified six comparative studies of AI-assisted diabetic-retinopathy pathways. We reconstructed all six 2×2 tables and reproduced the published result to the reported precision.
Across the six, recorded referral attendance was higher in the AI-enabled pathway: pooled risk ratio 1.89, with a 95% confidence interval from 1.18 to 3.03. The pooled absolute difference was 23.9 percentage points, from 12.8 to 35.1.
This is the strongest comparative evidence we found that AI-enabled pathways are associated with higher recorded referral attendance.
It is not a clean estimate of what the algorithm itself caused.
The table below reproduces the review’s published-synthesis extraction:
| Comparative study | AI-pathway attendance | Comparator attendance | Risk ratio |
|---|---|---|---|
| ACCESS, United States | 16/25, 64.0% | 18/83, 21.7% | 2.95 |
| DeepDR-LLM, China | 112/144, 77.8% | 90/154, 58.4% | 1.33 |
| EyeArt follow-up, United States | 51/92, 55.4% | 182/974, 18.7% | 2.97 |
| RAIDERS, Rwanda | 70/136, 51.5% | 55/139, 39.6% | 1.30 |
| STATUS, United States | 99/279, 35.5% | 14/117, 12.0% | 2.97 |
| Thailand digital platform | 115/129, 89.1% | 124/175, 70.9% | 1.26 |
STATUS also reported 12 internal visits among 103 referral-eligible patients in its nested AI–human salvage branch. The table follows the review’s AI-versus-historical-control comparison; the hybrid branch is not omitted from the evidence record or added as an independent cohort.
For ACCESS, the review uses 18/83 in the control arm. The primary analysis uses 18/82 after excluding one control participant enrolled in another screening study; our primary-endpoint sensitivity analysis uses 18/82.
ACCESS is also stage-asymmetric by design. The AI numerator is 16 of 25 participants with a disease-present result who completed a follow-up eye-care visit. The control numerator is 18 participants who completed an initial diabetic-eye examination after routine referral; all 18 were found not to have disease. That is a defensible end-to-end pathway contrast, but not a like-for-like referral-uptake comparison among screen-positive patients. A stage-matched control risk ratio cannot be recovered because the control arm did not identify an equivalent disease-positive denominator.
The studies disagree sharply. I², a measure of between-study inconsistency, was 91.9%. The 95% prediction interval—our estimate of what a new similar setting might show—ran from a risk ratio of 0.53 to 6.78, spanning a possible decrease through a very large increase.
We also repeated the analysis while omitting one study at a time. With the review’s published rows, one of six confidence intervals crossed 1; after correcting the ACCESS control denominator and substituting Thailand’s closer stage comparison, three of six did.
The comparison arms differ too. In ACCESS, the control group received a routine referral recommendation while the AI group received a point-of-care result. In RAIDERS, the randomized contrast was an immediate AI result against delayed human grading. Other studies add counselling, scheduling, reminders or new referral rules. Our coding found at least two workflow differences in every comparison and as many as four; unreported components may raise that count.
Remove ACCESS and EyeArt—the two comparisons whose controls received routine or universal referral advice—and align Thailand to the same care stage in both arms. The remaining DeepDR-LLM, RAIDERS, STATUS and Thailand rows produce a descriptive pooled risk ratio of 1.47, with a confidence interval from 0.80 to 2.71. That restriction removes two of the three largest individual risk ratios. The two randomized trials alone give a risk ratio of 1.90 with a 95% confidence interval from 0.01 to 341.83, far too wide to guide a stable conclusion.
The pooled average favors AI-enabled pathways. Whether that result will travel to a new setting is unresolved.
At least two rows do not compare the same clinical transition in both arms. ACCESS is one. Thailand is the clearest numerical mismatch.
In the implementation study, 115 of 129 AI-pathway true-positive referrals attended tertiary care. The review’s narrative compares that 89.1% with the primary paper’s stage-matched 17 of 22, or 77.3%, and reports p=.158. Yet its synthesis table uses 124 of 175, or 70.9%, for the manual period; that 124 appears to combine 116 intermediate confirmation visits and eight direct tertiary visits. The periods are observational and not directly comparable, but the review still pooled the less aligned of its own two manual-period numbers.
Correcting only that row barely changes the pooled point estimate: risk ratio 1.87, with a 95% confidence interval from 1.14 to 3.06. The result survives. The interpretation changes. The original rows do not all estimate the same clinical transition, and the corrected prediction interval remains wide, 0.49 to 7.07.
The six AI denominators in the review’s main synthesis table sum to 805. Its supplementary summary-of-findings table instead says “AI-assisted screening: 889,” 84 more. The manual-arm denominators reconcile exactly at 1,642, but neither the main paper nor its two supplements explains the extra 84 on the AI side.
That internal discrepancy remains unresolved in the public report. Neither 805 nor 889 is a shared screening denominator; both describe study-defined groups entering different comparisons.
On average, AI-enabled pathways recorded more referral attendance, but they tested whole workflows, not the algorithm alone. The review summarized the absolute difference as roughly one extra referral completion per four eligible people. That is not a screening-wide number needed to treat: eligibility varied across studies, and our prediction interval for the absolute difference ran from −1.6 to +49.5 percentage points, crossing zero.
A positive mean, a wide range of plausible settings
The forest separates confidence intervals from prediction intervals. The matrix keeps the surrounding workflow bundle visible.
AI-enabled pathways raised recorded attendance on average—not reliably everywhere
View chart data
| Comparison | Risk ratio | Limits | Interval type | Recorded attendance | Interval vs 1 | Design note |
|---|---|---|---|---|---|---|
| ACCESS | 2.95 | 1.78 to 4.88 | 95% confidence interval | 16/25 vs 18/83 | excludes 1 | Randomized controlled trial; low risk of bias; stage-asymmetric transition |
| DeepDR-LLM | 1.33 | 1.13 to 1.56 | 95% confidence interval | 112/144 vs 90/154 | excludes 1 | Sequential prospective comparison; moderate risk of bias |
| EyeArt follow-up | 2.97 | 2.37 to 3.72 | 95% confidence interval | 51/92 vs 182/974 | excludes 1 | Prospective cohort with historical comparator; serious risk of bias |
| RAIDERS | 1.30 | 1.00 to 1.69 | 95% confidence interval | 70/136 vs 55/139 | excludes 1 | Randomized controlled trial; low risk of bias |
| STATUS | 2.97 | 1.77 to 4.97 | 95% confidence interval | 99/279 vs 14/117 | excludes 1 | Historical workflow comparison; serious risk of bias |
| Thailand · published | 1.26 | 1.12 to 1.41 | 95% confidence interval | 115/129 vs 124/175 | excludes 1 | Alternating implementation periods; serious risk of bias |
| Thailand · same-stage | 1.15 | 0.91 to 1.46 | 95% confidence interval | 115/129 vs 17/22 | includes 1 | Alternative extraction of the same study; not a seventh comparison |
| Published pooled | 1.89 | 1.18 to 3.03 | 95% confidence interval, adjusted for six studies | 6 studies | excludes 1 | Random effects; between-study inconsistency (I²) 91.9% |
| Same-stage pooled | 1.87 | 1.14 to 3.06 | 95% confidence interval, adjusted for six studies | Thailand alternate | excludes 1 | Sensitivity analysis; I² 90.9% |
| Published new-setting range | 1.89 | 0.53 to 6.78 | 95% prediction interval | I² 91.9% | includes 1 | Expected range for a new setting; not a confidence interval |
| Same-stage new-setting range | 1.87 | 0.49 to 7.07 | 95% prediction interval | Thailand alternate | includes 1 | Prediction interval after stage-aligning the Thailand row |
- Both the published and same-stage pooled results are low-certainty, highly heterogeneous and pathway-specific. Their 95% prediction intervals include lower attendance in a new setting.
- The same-stage Thailand point is an alternative extraction of the same study, not a seventh independent comparison.
- ACCESS is stage-asymmetric by design: 16/25 is follow-up among AI disease-present patients, while 18/83 is an initial eye examination among the entire routinely referred control arm. It is not a like-for-like uptake comparison among screen-positive patients.
- The six rows use study-defined denominators and mixed designs. They estimate bundled pathways, not the classifier alone. The main-table AI denominators sum to 805; Supplementary Table 7 prints 889, an unexplained difference of 84.
- In the accessible table, confidence-interval limits are the lower and upper bounds around each estimate. Prediction intervals describe the wider range a new similar setting might show.
Every comparison changed more than the classifier
| Comparison | Instant result | Same-day option | Counselling | Scheduling | Reminders | Navigation / transport | Changed referral rule |
|---|---|---|---|---|---|---|---|
| ACCESS2 added; routine or universal referral | added in AI pathway | not documented in the AI pathway | documented in both pathways | not documented in the AI pathway | not documented in the AI pathway | not documented in the AI pathway | added in AI pathway |
| DeepDR-LLM3 added; manual grading workflow | added in AI pathway | not documented in the AI pathway | added in AI pathway | not documented in the AI pathway | not documented in the AI pathway | not documented in the AI pathway | added in AI pathway |
| EyeArt follow-up4 added; routine or universal referral | added in AI pathway | not documented in the AI pathway | added in AI pathway | added in AI pathway | not documented in the AI pathway | not documented in the AI pathway | added in AI pathway |
| RAIDERS2 added; manual grading workflow | added in AI pathway | added in AI pathway | documented in both pathways | not documented in the AI pathway | not documented in the AI pathway | documented in both pathways | not documented in the AI pathway |
| STATUS4 added; historical workflow | added in AI pathway | not documented in the AI pathway | added in AI pathway | added in AI pathway | not documented in the AI pathway | not documented in the AI pathway | added in AI pathway |
| Thailand3 added; manual grading workflow | added in AI pathway | not documented in the AI pathway | documented in both pathways | documented in both pathways | added in AI pathway | not documented in the AI pathway | added in AI pathway |
- A filled dot means the component was documented in the AI pathway and not documented as present in its comparator. It does not identify which component caused the attendance difference.
- Open dots keep shared counselling, scheduling or transport visible so those co-interventions are not silently credited to AI.
Three funnels, three missing links
No synthetic “average patient journey” can be made by splicing the best number from one study onto the next number from another. Three individual programs show why.
AEYE in an Israeli endocrinology clinic. In the 2026 real-world study, 256 people had an exam attempted, 245 received a definitive result and 76 were positive. The originating clinic documented 34 confirmatory examinations. Four of those 34 were judged to require treatment. The other 42 had no recorded internal confirmation but cannot be classified as having received no confirmatory care, because outside care was not captured. The paper does not report whether the four started or completed treatment, or what happened to vision.
Home AI screening in Boothgarh, India. In a three-arm pragmatic study, 200 people were enrolled in the home AI arm, 86 were referred and 15 reported by phone one month later that they had attended. Two participants had recorded diabetic-retinopathy treatment: one received laser, and one received laser plus an anti–vascular endothelial growth factor eye injection (supplementary Table 6). The paper does not report the intended treatment course, course completion or visual outcomes.
A linked publication from the same setting reported image gradability at eye level—224 of 362 analyzed eyes in the home/community AI arm—and that figure cannot be inserted as a patient-level step. The main adoption article and supplement also conflict on arm-level phone ascertainment, so we report attendance against the full referral group instead of selecting the more favorable contacted denominator.
A digital platform in Thailand. Here the chain extends farthest. Of 708 people screened in the AI period, 63 entered a separate visual-acuity referral branch. Among the remaining 645, AI flagged 201; specialist overread retained 129; 115 attended tertiary care; 48 had vision-threatening diabetic retinopathy; and 18 received laser or an eye injection.
The side branches matter. Of the 63 visual-acuity referrals, 52 attended. Overread rejected 72 AI positives, although 17 still presented, and found five false negatives, four of whom attended. The 67 ungradable attenders included two cataract treatments, which we do not count as diabetic-retinopathy treatment. No completed treatment course or vision outcome was reported.
The Thailand study documents high recorded attendance within a redesigned pathway; it does not isolate an AI effect or establish superiority over the manual period. The program combined AI, specialist overread, reminders and active referral cancellation. Credit belongs to the pathway.
Three programs, three separate missing links
The cascades stay separate because their gates, capture domains and endpoints differ. Side branches remain visible rather than being converted into attrition.
AEYE real-world clinic: outside follow-up breaks the chain
View chart data
| Stage | What happens | Note |
|---|---|---|
| Exam attempted · 256 people | ||
| Definitive result · 245 | 95.7% of attempts | |
| Positive result · 76 | 31.0% of definitive results; 29.7% of attempts | |
| (observed internally) Internal confirmation · 34 | 44.7% of positive results | |
| (outside care unobserved) No internal confirmation record · 42 | Cannot be classified as no confirmatory care | |
| Treatment judged necessary · 4 | 11.8% of internally examined patients | |
| Initiation · completion · vision | Not reported |
- The 42 people without an internal confirmation record may include outside care. They are unobserved, not verified nonattenders.
- Treatment judged necessary is not evidence that treatment began or was completed.
Boothgarh home AI: treatment status recorded for two; completion unknown
View chart data
| Stage | What happens | Note |
|---|---|---|
| Home AI arm · 200 people | ||
| Referred · 86 | 43.0% of the arm | |
| Recorded attendance at one month · 15 | 17.4% of referrals | The remaining 71 have no recorded visit; the study does not verify all as nonattenders. |
| Participants with recorded diabetic-retinopathy treatment · 2 | One retinal laser; one retinal laser plus a drug injection into the eye (anti-VEGF) | Among 11 ungradable attenders, the supplement lists four as 'cat surgery,' three as 'cat Sx appointment' and four as no treatment. Completed cataract surgery is not clearly established; these categories are not counted as diabetic-retinopathy treatment. |
| Completed diabetic-retinopathy episode · vision outcome | Not reported |
- Arm-level phone-ascertainment totals conflict between the main text and supplement, so the chart does not calculate attendance only among people reached.
- The linked 224/362 gradability count is eye-level and is not inserted into this patient-level funnel.
Thailand: the longest reported chain still has side branches
View chart data
| Stage | What happens | Note |
|---|---|---|
| Screened in AI period · 708 | ||
| (VA ≤20/70) Separate visual-acuity referral · 63 | 52 later presented at tertiary care | |
| (AI workflow denominator) Entered AI-image comparison · 645 | The 201 AI positives use this denominator—not 708 | |
| AI referral-positive · 201 | 31.2% of 645 | Of 444 AI negatives, overread identified five false negatives; four attended after contact. |
| (true-positive referrals) Over-reader agreed · 129 | 64.2% of AI positives | |
| (false-positive branch) Over-reader disagreed · 72 | 17 still presented | 49 were reached to reverse referral (2 attended); 23 were not reached (15 attended). |
| True-positive referrals attending tertiary care · 115 | 89.1% of 129 | |
| Vision-threatening diabetic retinopathy confirmed · 48 | 45 diabetic macular edema; 3 severe retinopathy | |
| Retinal laser or eye injection received · 18 | 25 monitored; 5 referred onward | |
| Completed treatment episode · vision outcome | Not reported |
- The main spine follows over-reader-agreed true-positive referrals. It does not absorb the 52 VA-branch attenders, 17 false-positive presentations or four false-negative attenders into the 115 denominator.
- A separate 67-person ungradable branch included two cataract treatments; those are not diabetic-retinopathy treatment events.
- The pathway combined AI, specialist overread, reminders and referral reversal. The chart cannot isolate an algorithm-only effect.
A million screens, outcomes uncounted
The language of deployment moves more freely than the evidence.
In March, Google reported more than one million screenings through clinical partnerships across India, Thailand and Australia. The company says a diagnosis can arrive “in as little as two minutes” and “could potentially save their sight.” The sentence is appropriately hedged. Ground Truth did not independently verify the scale, and the page does not say whether each screening represents a unique person. The linked public studies do not provide one program-wide count moving from result to confirmation, completed treatment and vision.
Digital Diagnostics uses “Prevent blindness through early disease detection” as product language for LumineticsCore. The linked STATUS implementation study documents nine patients who “received treatment … at the first follow-up encounter.” That is further down the chain than diagnostic accuracy alone. It is still not a completed treatment plan or a longitudinal vision outcome.
Orbis says patients given real-time AI diagnoses “are more likely to seek treatment.” The underlying RAIDERS randomized trial measured presentation for referral services within 30 days; it did not measure treatment initiation.
Other current scale statements have the same structural limit. Eyenuk has reported more than 230,000 patients screened.
Remidio uses the same 16-million figure for different stages: its product page displays “16M+ Patients Screened,” while an NITI Frontier Tech profile submitted under the Frontier Voices programme describes “16M+ people” risk-triaged and separately lists “1.2M+ retinal scans” analyzed. The profile says those scans generated epidemiological insights, including glaucoma-prevalence estimates.
India’s MadhuNetrAI announcement says 7,100 patients are “benefiting.” None of those public numbers, as currently linked, supplies a complete program-wide treatment and vision denominator.
These statements may each be accurate at their own stage, but they are not interchangeable. Screening volume is a screening-volume claim. Referral attendance is a referral claim. Treatment initiation is a treatment claim. Preserved sight is an outcome claim.
Each requires its own denominator.
The fair case for deploying before the final outcome
A screening program does not need to wait for a definitive population-level blindness trial before using an accurate, regulated test.
The clinical logic is established: diabetic retinopathy can be asymptomatic, timely detection matters, and effective laser and injection treatments already exist. Vision outcomes take larger samples and longer follow-up than diagnostic studies. External care is difficult to observe. A same-day answer can remove delays even when the algorithm is not the only active ingredient.
The ACCESS trial and RAIDERS trial support that operational case. Immediate, point-of-care pathways can improve recorded follow-through in some settings. A product can be useful before every downstream question is answered.
But the burden changes with the claim. “Returns a diagnostic result at the point of care” can be supported by an accuracy and technical-yield study. “Improves referral” requires a credible comparator, aligned care stages and comparable follow-up capture and timing. “Prevents blindness” requires comparative longitudinal evidence that the program reduced vision loss or blindness, with treatment and follow-up denominators.
Programs large enough to report hundreds of thousands or millions of screens are also large enough to make the missing denominators consequential. If follow-up occurs outside the screening system, that is a reason to build linkage or sampling—not a reason to call the outcome known.
What proof would look like
The next retinal-AI study does not need another isolated sensitivity headline. It needs a linked ledger of unique people.
For every person with an image attempted, report:
- whether repeat capture or dilation was needed;
- whether the system returned a definitive result;
- whether the result was positive or ungradable, and the rule for each;
- whether referral was recommended;
- whether a confirmatory examination occurred, including outside the originating system;
- whether referable or vision-threatening disease was confirmed;
- whether treatment was indicated;
- whether treatment began;
- whether the intended treatment episode was completed; and
- visual acuity or another longitudinal vision outcome at a stated time.
Every transition needs a numerator, its immediate denominator and a time window. Repeated cameras and repeated visits need cohort identifiers so one person remains one person. Unknown follow-up should stay unknown. Workflow components—same-day results, counselling, scheduling, reminders, navigation, transport and specialist overread—should be recorded rather than credited silently to the classifier.
This is not an impossible standard. It is the ordinary accounting required to move from a test to a health outcome.
The bottom line
Can the AI eye exam read the retina? In many settings, yes. The diagnostic evidence is substantial, and our reconstructed AEYE matrices reproduce the reported sensitivity and specificity.
Can an AI-enabled pathway help more people reach eye care? Sometimes. The six-study pooled result is positive, but it is highly heterogeneous, denominator-sensitive and inseparable from the workflow around the model.
Can the current public evidence show that these programs complete treatment and preserve sight at scale? Not yet. Across 17 selected pathway rows representing 14,261 people, two of those rows independently report 10 treatment indications. Six report 56 treatment or management events, but only 33 have a named modality or a clearly documented first-treatment event. No selected row reports completion of the intended treatment course or a denominator for longitudinal vision outcome.
Zero measured vision outcomes is an evidence gap, not evidence of zero benefit.
The public record shows where the patient entered the funnel. It rarely shows where the patient emerged.
How this analysis was built
Ground Truth searched public evidence through 30 August 2026 and structured 29 unique primary or implementation studies into 41 arm or cohort rows across 14 countries. Seventeen studies entered the pathway-outcome census; six paired comparisons entered the referral synthesis. We separately mapped 15 current public claims to the furthest endpoint reported in their linked evidence. This was a citation-seeded evidence census, not a PRISMA-complete systematic review.
The protocol was frozen before the full source census but after several seed findings were known. It is not preregistered. We preserved unique-person lineage, patient-versus-eye units, image-attempt and definitive-result denominators, follow-up ascertainment, internal-versus-external capture, and treatment indication, initiation and completion as separate fields.
The review describes a Mantel–Haenszel analysis in R’s meta package. In that implementation, the random-effects summary uses inverse-variance weighting; with REML estimation and Hartung–Knapp confidence intervals, we reproduced the published risk ratio, risk difference and heterogeneity to the reported precision. We then ran stage-alignment, randomized-only, comparator-rule, risk-of-bias and leave-one-out analyses.
Three separately scoped research passes were reconciled against primary sources. The executable package validates 58 source records, 41 study rows, 15 claims, 36 aggregate facts and three diagnostic 2×2 tables. Thirteen gold tests pass. All 58 locally archived source files pass byte, hash and format checks. The AEYE pivotal Figure 2 is archived as its own hashed primary-source asset; the STATUS patient-flow figure is preserved inside a hashed supplementary ZIP, and the referral review’s DOCX and PDF supplements are separately archived and hashed.
The reusable data are published under CC BY 4.0: download the 17-study patient-count ledger, comparative referral effects, sensitivity analyses, diagnostic reconstructions, claim-endpoint map and source manifest.
This analysis is based entirely on the public record; no company or study-author outreach is part of the publication workflow. Corrections can be sent to corrections@groundtruth.health and will be logged publicly when they change the record.
Key sources
- AEYE pivotal synthesis: Frontiers in Digital Health, 25 August 2026
- AEYE real-world pathway: British Journal of Ophthalmology / PubMed
- Regulator-approved diagnostic-system review: npj Digital Medicine
- Referral-uptake systematic review and meta-analysis: npj Digital Medicine
- UK National Screening Committee external review: Automated grading in diabetic eye screening, 2026
- Thailand national diagnostic evaluation: The Lancet Digital Health / PubMed
- Thailand digital-platform implementation: Ophthalmology and Therapy / PMC
- ACCESS randomized trial: Nature Communications / PMC
- RAIDERS randomized trial: Ophthalmology Science / PMC
- STATUS implementation study: Clinical Ophthalmology / PMC
- EyeArt follow-up study: Ophthalmology Retina / PubMed
- Boothgarh pragmatic trial: Archives of Public Health / PMC
- Moga implementation study: JMIR Medical Informatics / PMC
- Australian implementation study: Scientific Reports / PMC
- Rural Maine EyeArt study: Journal of General Internal Medicine / PMC
Disclosures & provenance
- Published
- 31 Aug 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Data
- Download the dataset · released under CC BY 4.0
- Corrections
- None to date. Corrections log → · Challenge this analysis