Investigation · Diagnostic AI
“Predicts risk of more than 300 diseases”: what the Aladynoulli paper reports, and what peer review narrowed
Harvard Medical School says a new AI tool predicts the risk of more than 300 diseases from existing patient data. The Nature paper behind the release is a careful piece of work: it outputs a risk for 348 PheCode disease categories, publishes an accuracy figure for 28 of them, tested prediction in one biobank, and — after a referee objected — removed its general claims of “interpretability.” The release describes the model’s signatures as “interpretable,” calls it “the first known” tool of its kind, and attaches “high accuracy” to a colorectal-cancer use the paper never discusses.
On 17 August 2026, Harvard Medical School published a news item headlined “New AI Tool Predicts Risk of More Than 300 Diseases With Existing Patient Data.” Its opening line describes “the first known machine-learning-powered tool that can predict patient risk of hundreds of diseases based solely on patient health records and genetic profiles.” The item is adapted from a Dana-Farber Cancer Institute release of 15 July, the day the paper appeared in Nature.
The paper is real and open access: “A Bayesian framework for longitudinal EHR and genetic discovery” by Sarah Urbut and colleagues at Mass General Brigham, Dana-Farber and the Harvard Chan School, with Alexander Gusev, Pradeep Natarajan and Giovanni Parmigiani as senior authors. It ships a 48-page supplement and an 88-page peer-review file: three referees, two rounds, both author responses.
We rate the release Overstated (2/5). The numbers it uses are real and traceable. The model, called ALADYNOULLI, does output a risk for 348 PheCode disease categories; it was fitted to three biobanks totalling 683,571 people; at ten years it does beat three established heart-disease scores on the same 400,000 UK Biobank participants, with confidence intervals that do not overlap. But the measurement sits well inside the words. An accuracy figure is published for 28 of the 348 diseases; prediction was tested in one biobank; every person it was tested on had genotype data and 36 computed polygenic risk scores — the release’s opening line discloses the genetic profiles, its headline’s “existing patient data” does not, and no genetics-free performance result is reported; the “first known” sentence appears nowhere in the paper, which instead benchmarks against an earlier multi-disease model; and the “high accuracy” colorectal-cancer scenario has no threshold, no positive predictive value and no text in the paper at all. The release carries one caveat the paper would endorse — “At this time, the tool has only been tested on past patient data” — and that sentence is the reason this is not rated lower.
The most telling document is the peer-review file. In it, Referee #2 wrote that “some of the claims oversell the package to its detriment.” The authors agreed: “we have removed claims of general ‘interpretability’,” they replied on 3 June 2026, and “every instance of ‘leakage-free’ has been removed from the manuscript.” Six weeks later the institutional release quoted a senior author saying the model’s “curated signatures capture the underlying biology in an interpretable way.”
Data table
| Stage | What happens | Note |
|---|---|---|
| Three biobanks, 683,571 participants | UK Biobank 427,239 · Mass General Brigham 48,069 · All of Us 208,263. ICD-10 codes mapped to 348 PheCode disease categories (each with ≥1,000 UK Biobank occurrences). Genetic and demographic covariates: 36 polygenic risk scores, sex and 10 genetic principal components. | |
| Signature discovery: 21 latent signatures (20 disease + 1 low-incidence) | A Bayesian generative model is fitted in each cohort; disease-to-signature membership is compared across cohorts (median composition preservation 0.80); signature loadings feed a GWAS (151 loci) and rare-variant and carrier analyses. | |
| Prediction evaluation: UK Biobank only | 400,000 participants in 40 batches of 10,000, each held out in turn. One-year risk predicted at enrolment + 0 … 9 years; the “ten-year” figure is the one-year risk at recruitment scored against ten-year outcomes. Enrolment 2006–2010; median follow-up 14.4 years. | |
| Discrimination reported for 28 diseases | Supplementary Table 14: one-year AUC at enrolment and ten-year AUC with 95% CIs, plus a one-year median across enrolment and years 1–9 without intervals. Calibration is reported once, pooled across all 348 PheCodes. | |
| Clinical risk scores on the same 400,000 participants | Ten-year ASCVD: ALADYNOULLI 0.733 · QRISK3 0.702 · PCE 0.683 · PREVENT 0.667 (34,705 events). Breast cancer versus Gail: one-year 0.782 vs 0.549; ten-year 0.554 vs 0.540. | |
| Peer review at Nature: three referees, two rounds (Aug 2025 – Jun 2026) | After Referee #2 objected, the authors removed general “interpretability” claims and every “leakage-free”; the limitations section grew from five items to seven. The published title is “A Bayesian framework for longitudinal EHR and genetic discovery.” | |
| 15 Jul 2026: paper and Dana-Farber release · 17 Aug 2026: HMS release | HMS: “New AI Tool Predicts Risk of More Than 300 Diseases With Existing Patient Data” — “the first known machine-learning-powered tool” — “our curated signatures capture the underlying biology in an interpretable way” — and a caveat: “At this time, the tool has only been tested on past patient data.” |
- This diagram introduces no new numbers: every figure on it is printed in the paper, its supplement or the peer-review file, and is discussed on this page with its source.
- The peer-review box describes what the published reviewer reports and author responses say. Who initiated the change of title is not recorded in the file, and we do not say.
- The release box quotes the HMS page as published on 17 August 2026. The Dana-Farber release of 15 July carries the same Parmigiani passage without Harvard's bracketed insertion, calls the model “the first” (Harvard: “the first known”), calls the colorectal result “very high accuracy,” and does not carry Harvard's retrospective-data caveat.
The claim, and the caveat it carries
The Harvard item runs to under a thousand words. Its “At a glance” box states the model is “the first known machine-learning-powered tool” of its type; the body calls it “a machine-learning algorithm,” “an artificial intelligence-based model” and “the first to make multi-disease predictions based only on routinely collected electronic health record data and knowledge about the patient’s genetic risks of disease.” Its evidence paragraph reads, in full: “The model was trained and validated using three large biobanks including a total of over 683,000 patient records. The model provided more accurate 10-year predictions than three existing cardiovascular risk models — PCE, QRISK3, and PREVENT. It also outperformed the GAIL breast cancer risk model for one-year breast cancer predictions.” The next paragraph adds that the model “can also predict which patients will develop colorectal cancer in the coming year with high accuracy,” and that flagging such a patient “could potentially help a primary care physician refer a patient for a colonoscopy, even if that patient is not yet eligible for screening.”
It also says, plainly: “At this time, the tool has only been tested on past patient data. Future work is needed to assess it in clinical settings, as a research tool, and as a way to improve the design of clinical trials.” That is accurate, and the paper’s seven stated limitations say the same at greater length. The Dana-Farber original uses much of the same material but differs materially: it calls the model “the first” where Harvard’s adaptation says “the first known,” calls the colorectal result “very high accuracy,” and does not include Harvard’s “only been tested on past patient data” paragraph. The Boston Globe, interviewing the senior authors on publication day, quoted Parmigiani saying the technology is “ready for prime time,” and in the next sentence reported that the team could not say when it might be deployed.
What the paper is: a discovery framework, with prediction as one of its applications
ALADYNOULLI is a Bayesian generative model. It takes a person’s diagnosis history — ICD-10 codes mapped onto 348 “PheCodes,” each with at least a thousand occurrences in the UK Biobank — together with age, sex, ten genetic principal components and 36 polygenic risk scores, and it explains that history as a mixture of 21 latent “signatures”: 20 disease signatures and one low-incidence signature, each a set of diseases that tend to travel together and a curve describing how their probabilities change with age. The signatures were fitted separately in three cohorts — UK Biobank (427,239 people), Mass General Brigham (48,069) and All of Us (208,263) — and the disease-to-signature memberships matched across them with a median “composition preservation” of 0.80. The signature loadings then drove a genome-wide association study that found 151 loci, rare-variant analyses that picked out LDLR, TTN and BRCA2, and checks that carriers of familial hypercholesterolaemia load onto the cardiovascular signature and carriers of clonal haematopoiesis onto the inflammatory one.
That is the bulk of the paper, and it is what the three cohorts validate. Prediction is one of five results sections, one of five main figures, and roughly a fifth of the results text. The abstract gives it one sentence, its last: ALADYNOULLI “outperforms Pooled Cohort Equation (PCE), PREVENT and Gail at 1-year and 10-year horizons; disease-level (PheCode) predictions complement code-level foundation models such as Delphi-2M.” The title on the published version is about discovery. The title on the manuscript the authors first resubmitted was “ALADYNOULLI: A Bayesian approach to disease progression modeling for genomic discovery and clinical prediction”; the words “clinical prediction” were gone by the second revision. The peer-review file does not record who proposed that change.
The news item is a prediction-tool story built from a paper in which prediction is the last application described.
348 diseases modelled, 28 evaluated, one biobank tested
Three numbers set the scale of the gap. The first is 348 — the number of PheCodes for which the model produces a risk. That is the release’s “more than 300.” The second is 28 — the number of diseases for which the paper publishes a discrimination figure, an AUC, in Supplementary Table 14. The third is one — the number of biobanks in which prediction was evaluated. The Methods are explicit: the genetic and performance results are based on the UK Biobank; the model was also trained in Mass General Brigham and All of Us “to establish consistency,” and “No prediction tasks were performed on these two cohorts.” The release’s sentence that the model was “trained and validated using three large biobanks” is true of the signatures. It is not true of the predictions. (The paper’s Reporting Summary, a form completed for the journal, says performance “was assessed in these biobanks separately”; no such result appears in the paper or its supplement.)
Data table
| Domain | 10-year AUC (95% CI) | 95% CI | CI vs 0.5 | Marker | 1-year AUC at enrolment |
|---|---|---|---|---|---|
| ASCVD | 0.733 | 0.73 to 0.736 | above | filled | 0.881 |
| Parkinson's disease | 0.724 | 0.715 to 0.735 | above | filled | 0.809 |
| Bladder cancer | 0.708 | 0.698 to 0.716 | above | filled | 0.825 |
| Chronic kidney disease | 0.708 | 0.703 to 0.712 | above | filled | 0.651 |
| Atrial fibrillation | 0.707 | 0.703 to 0.711 | above | filled | 0.797 |
| Heart failure | 0.701 | 0.696 to 0.707 | above | filled | 0.769 |
| Prostate cancer | 0.687 | 0.682 to 0.693 | above | filled | 0.831 |
| Osteoporosis | 0.681 | 0.676 to 0.686 | above | filled | 0.756 |
| Stroke | 0.681 | 0.675 to 0.688 | above | filled | 0.653 |
| All cancers | 0.674 | 0.67 to 0.677 | above | filled | 0.753 |
| Lung cancer | 0.669 | 0.663 to 0.677 | above | filled | 0.699 |
| COPD | 0.658 | 0.655 to 0.662 | above | filled | 0.736 |
| Diabetes | 0.651 | 0.648 to 0.654 | above | filled | 0.741 |
| Colorectal cancer | 0.648 | 0.641 to 0.655 | above | filled | 0.825 |
| Pneumonia | 0.644 | 0.639 to 0.649 | above | filled | 0.634 |
| Secondary cancer | 0.61 | 0.605 to 0.616 | above | filled | 0.600 |
| Rheumatoid arthritis | 0.608 | 0.601 to 0.614 | above | filled | 0.749 |
| Thyroid disorders | 0.594 | 0.59 to 0.598 | above | filled | 0.678 |
| Multiple sclerosis | 0.591 | 0.571 to 0.609 | above | filled | 0.840 |
| Anemia | 0.588 | 0.584 to 0.592 | above | filled | 0.648 |
| Ulcerative colitis | 0.583 | 0.57 to 0.599 | above | filled | 0.816 |
| Crohn's disease | 0.58 | 0.563 to 0.598 | above | filled | 0.896 |
| Breast cancer | 0.554 | 0.548 to 0.56 | above | filled | 0.782 |
| Psoriasis | 0.546 | 0.532 to 0.558 | above | filled | 0.607 |
| Asthma | 0.529 | 0.525 to 0.533 | above | filled | 0.690 |
| Anxiety | 0.514 | 0.509 to 0.52 | above | filled | 0.604 |
| Bipolar disorder | 0.492 | 0.472 to 0.513 | crosses | hollow (CI crosses 0.5) | 0.758 |
| Depression | 0.484 | 0.478 to 0.488 | below | hollow (CI below 0.5) | 0.616 |
- The model outputs risks for 348 PheCodes; a discrimination figure (AUC) is published for these 28. The release's “more than 300 diseases” is the first number; this chart is the second. Calibration is reported once, pooled across all 348.
- All prediction results are from the UK Biobank only — 400,000 participants in 40 batches of 10,000, each held out in turn, enrolled 2006–2010 at ages the main text gives as 40–70 (median 54) and Supplementary Fig. 14 as 37–73 (median 59). No prediction task was run in Mass General Brigham or All of Us; those cohorts were used to check that the disease signatures replicate.
- “Ten-year” here is the model's one-year risk at recruitment, ranked against whether the disease occurred within ten years (Methods). It is a discrimination number; no ten-year calibration is reported.
- Eleven of the 28 ten-year AUCs are at or below 0.60. Depression's interval lies wholly below 0.5 (0.484, CI 0.478–0.488); bipolar disorder's point estimate is below 0.5 with a CI that crosses it (both hollow markers); anxiety is just above (0.514). The one-year figures in the note column are higher for 24 of the 28 — CKD, pneumonia, secondary cancer and stroke are the exceptions — so prediction is generally strongest for the year immediately ahead.
- CIs are from 100 bootstrap resamples of the held-out test sets; the table's “CI vs 0.5” column records whether each interval lies above, crosses, or lies below 0.5 — no hypothesis test is implied. The spectral-clustering initialisation of the model's signature–disease starting values was run once on the entire dataset and reused in every batch (Methods); the authors do not quantify any effect of this on held-out AUCs, and we do not claim one.
The 28 figures themselves are mixed, and the paper does not hide that. At one year from enrolment, ALADYNOULLI discriminates well for several conditions — 0.881 for atherosclerotic cardiovascular disease, 0.896 for Crohn’s disease, 0.840 for multiple sclerosis, 0.825 for colorectal cancer. The ten-year figures are lower for 24 of the 28 — CKD, pneumonia, secondary cancer and stroke are the exceptions — as the authors note they would be: 0.733 for ASCVD, 0.651 for diabetes, 0.554 for breast cancer. Eleven of the 28 are at or below 0.60. Depression’s interval lies wholly below 0.5 (0.484, 0.478–0.488); bipolar disorder’s point estimate is below 0.5 with an interval that crosses it; anxiety is 0.514. They are in the supplement with confidence intervals, to the authors’ credit. They are not in the release — and “predicts risk of more than 300 diseases” does not prepare a reader for them.
Calibration — whether the predicted risks match the rates that occurred — is reported once, pooled across all 348 diseases and 722 million person-time observations: mean predicted rate 0.000555 against mean observed 0.000545. That is calibration in the large for the whole model. No disease-by-disease calibration is published, and for a tool whose proposed use is flagging individuals, per-disease calibration is the number that would matter.
The comparisons that are fair, and the figure that mixes horizons
Start with the result that holds. PCE, PREVENT and QRISK3 are ten-year cardiovascular risk scores. Scored on ten-year outcomes in the same 400,000 UK Biobank participants, with the same 34,705 events, they reach AUCs of 0.683, 0.667 and 0.702. ALADYNOULLI reaches 0.733 (95% CI 0.730–0.736). The intervals do not overlap, and the margins — +0.031 over QRISK3, +0.050 over PCE, +0.066 over PREVENT — are real. The release’s sentence about “more accurate 10-year predictions than three existing cardiovascular risk models” is supported. Two details travel with it. The model’s “ten-year” number is its one-year risk at recruitment, ranked against whether the disease occurred within ten years, whereas the clinical scores output a ten-year absolute risk; that is a discrimination contest, and no ten-year calibration comparison is published. And the paper is of two minds about how the comparison was tested: the Figure 5 caption says the methods were compared “via overlap of their 95% bootstrap confidence intervals; no separate hypothesis test was applied,” while the Supplementary Figure 17 caption describes a two-sided bootstrap-difference test. Neither place prints a P value. The inconsistency is the point; the absence of a P value on its own is not a defect.
Data table
| Domain | AUC (95% CI) | 95% CI | Horizon | Marker | Horizon · sample |
|---|---|---|---|---|---|
| ASCVD · ALADYNOULLI | 0.733 | 0.73 to 0.736 | 10-year | filled | 10-yr · 400,000 · 34,705 events |
| ASCVD · QRISK3 | 0.702 | 0.699 to 0.705 | 10-year | filled | 10-yr · same |
| ASCVD · PCE | 0.683 | 0.681 to 0.685 | 10-year | filled | 10-yr · same |
| ASCVD · PREVENT | 0.667 | 0.665 to 0.669 | 10-year | filled | 10-yr · same |
| Breast cancer · ALADYNOULLI | 0.782 | 0.759 to 0.81 | 1-year | filled | 1-yr · women |
| Breast cancer · Gail | 0.549 | 0.529 to 0.567 | 1-year | filled | 1-yr · women |
| Breast cancer · ALADYNOULLI | 0.554 | 0.548 to 0.56 | 10-year | filled | 10-yr · women |
| Breast cancer · Gail | 0.54 | 0.534 to 0.545 | 10-year | filled | 10-yr · women |
- This is the fair comparison, and it favours the model: PCE, PREVENT and QRISK3 are ten-year scores, scored here on ten-year outcomes in the same 400,000 participants with the same 34,705 ASCVD events, and the confidence intervals do not overlap. The margins are +0.031 (QRISK3), +0.050 (PCE) and +0.066 (PREVENT).
- The paper's headline figure is a different comparison. Fig. 5b overlays ALADYNOULLI's one-year ROC curves (year 0 AUC 0.881) on PCE and PREVENT curves labelled “10 years (enrolment)” (0.678 and 0.653); Supplementary Fig. 15 scores the same two ten-year instruments on one-year outcomes (0.642 and 0.596). The mismatch is labelled in the legend and not discussed as a limitation.
- ALADYNOULLI's “ten-year” value is its one-year risk at recruitment ranked against ten-year outcomes; the clinical scores output ten-year absolute risk. No ten-year calibration comparison is published. The Fig. 5 caption says methods were compared by overlap of bootstrap CIs with no separate hypothesis test; the Supplementary Fig. 17 caption describes a two-sided bootstrap-difference test. No P value appears in either place.
- Breast cancer at ten years: the authors' response to reviewers calls 0.551 versus 0.540 “comparable”; Supplementary Table 14 gives 0.554 for the same quantity. The release cites only the one-year Gail result. Gail was computed “using the reported family history data”; which of the model's other inputs were available is not stated, and no code for PCE, PREVENT, QRISK3 or Gail is in the public repository.
- All four clinical scores were applied to UK Biobank volunteers, a healthier-than-average population; the authors address selection with inverse probability weighting for the discovery analyses.
The paper’s headline figure is a different comparison. Figure 5b overlays ALADYNOULLI’s one-year ROC curves — year-0 AUC 0.881, years 1–9 between 0.848 and 0.902 — on PCE and PREVENT curves labelled “10 years (enrolment),” at 0.678 and 0.653. Supplementary Figure 15 goes further and scores those two ten-year instruments on one-year outcomes, where they fall to 0.642 and 0.596 — and labels ALADYNOULLI’s own year-0 value as 0.870 there, against the 0.881 shown in Figure 5b and Table S14. The mismatch of horizons is visible in the legend and is not discussed or justified as a limitation; it is the version of the comparison a reader of the figure will remember, and the one a later Substack interview repeated as “0.89 versus 0.68”; the Boston Globe repeated only the general claim of outperformance. Two smaller details. The main text reports the fair ten-year values of 0.683 and 0.667 with a pointer to “Fig. 5d,” which is the calibration plot; they are in Supplementary Figure 17 and Table 16. And QRISK3 — the comparator the release names — is never given a number in the main text or the abstract; its 0.702 is in the supplement and the figure legend, and the score’s own paper is not cited.
Breast cancer is the same shape at sharper angle. At one year, ALADYNOULLI reaches 0.782 against 0.549 for the Gail model, a difference of +0.233 that the main text reports and the release repeats. At ten years, the longer-horizon comparison the authors reported, the difference is +0.011 — 0.551 against 0.540 in Supplementary Figure 17b, or 0.554 in Table 14 for the same quantity — a result the authors themselves called “comparable” in their response to reviewers. The abstract says the model outperforms Gail “at 1-year and 10-year horizons.” The release limits itself to the one-year result, which is the more defensible choice.
What the clinical scores were computed from is not described. The paper says Gail was run “using the reported family history data”; for PCE, PREVENT and QRISK3 it gives no input list, no missing-data handling, and — in the public code repository — no implementation. The comparison notebooks load pre-computed score files. That leaves a reader unable to check whether the scores were run as their authors intended.
“High accuracy” for colorectal cancer is an AUC, with no threshold and no referral analysis
The release says the model “can also predict which patients will develop colorectal cancer in the coming year with high accuracy,” and suggests such a flag could prompt a colonoscopy referral before screening age. Colorectal cancer is not discussed in the paper’s prose. It is a row in three supplementary tables and a label in two figures. The words “colonoscopy” and “referral” do not occur in the paper, the supplement or the peer-review file.
Data table
| Domain | AUC (95% CI) | 95% CI | Horizon | Marker | Horizon · group |
|---|---|---|---|---|---|
| All ages · ALADYNOULLI | 0.825 | 0.791 to 0.857 | 1-year | filled | 1-yr at enrolment |
| Age 39–50 · ALADYNOULLI | 0.729 | 0.589 to 0.8 | 1-year | filled | 1-yr · youngest reported stratum; overlaps but does not equal the release’s not-yet-eligible population |
| Age 50–60 · ALADYNOULLI | 0.837 | 0.811 to 0.867 | 1-year | filled | 1-yr |
| Age 60–72 · ALADYNOULLI | 0.907 | 0.891 to 0.927 | 1-year | filled | 1-yr |
| All ages · ALADYNOULLI | 0.648 | 0.641 to 0.655 | 10-year | filled | 10-yr at enrolment |
- Colorectal cancer is not discussed anywhere in the paper's prose; it appears as a row in three supplementary tables and as a label in two figures. The words “colonoscopy,” “referral” and “screening” do not occur in the paper or its supplement in connection with it. The referral scenario in the Dana-Farber and HMS releases is the institutions' framing.
- An AUC is a ranking statistic. It says nothing about what fraction of people flagged at a given threshold would go on to develop the cancer, and a positive predictive value cannot be inferred from an AUC and a base rate without a threshold. No threshold, sensitivity or positive predictive value is published for any disease. Over the ten-year rolling analysis in one 10,000-participant batch there were 105 colorectal events (1.1%).
- The youngest reported stratum (39–50) overlaps but does not equal the release's not-yet-eligible population; it has the lowest point estimate and the widest interval (0.589–0.800). The one-year figures are highest for the oldest group.
- The Cox baseline (age as timescale, sex and family history) reaches 0.521 at ten years in Supplementary Table 16; no confidence interval is published for it, so it is given here in text only.
- The one-year median across years 0–9 is 0.848 (no CI published). The ten-year figure is 0.648.
The numbers that exist are these. One-year AUC at enrolment: 0.825 (0.791–0.857). Median one-year AUC across the ten annual timepoints, enrolment and years 1–9: 0.848. Ten-year: 0.648. By age at enrolment, the one-year figure is 0.729 (0.589–0.800) for people aged 39–50, 0.837 for 50–60 and 0.907 for 60–72 — so the youngest reported stratum, which overlaps but does not equal the release’s not-yet-eligible population, is the one where the model is weakest and the interval widest. An AUC is a ranking statistic. It says how often a person who will develop the disease is ranked above one who won’t; it says nothing about how many of the people flagged at any given threshold will actually develop it. No threshold, sensitivity or positive predictive value is published for colorectal cancer or any other disease. In the ten-year rolling analysis on one 10,000-person batch there were 105 colorectal events, 1.1%. At that base rate a positive predictive value cannot be inferred from an AUC; the referral scenario needs a prespecified threshold and the positive predictive value that goes with it — the same arithmetic we set out for a “99% accurate” cancer test.
What the peer-review file shows the authors narrowed — and how the release described the model
Nature published the reviewer reports and author responses alongside the paper. They show a strong manuscript being made more careful. Referee #1 called the first version “mostly descriptive by nature” and asked for comparisons with “well-documented clinical risk scores” — the PCE, PREVENT, QRISK3 and Gail analyses were added in response. Referee #3 called the method “novel and impressive,” asked for selection-bias weighting, washout analyses and a comparison with the Delphi-2M transformer, and after the revision wrote, “I think this is a really strong paper.” Referee #2 pressed on two words.
“I must push back on the continuing use of phrases like ‘clinically interpretable’ and ‘biologically interpretable’ … replication is extremely important, but does not speak to biological or clinical interpretability; non-interpretable results can be replicated!”
Referee #2, second-round report, Peer Review File
“I continue to object to the characterization of this as a ‘leakage-free’ validation. Moreover, I believe there should be an explicit acknowledgement that this work treats the first date of a code as the true diagnosis date … descriptions of ‘leakage free’ should be removed from the paper since this is simply not possibly given the levels of uncertainty and heterogeneous error for diagnosis date.”
Referee #2, second-round report
“Despite these reservations, I think this is an interesting method that is well-execute on a number of levels. However, the above limitations are still under-described and some of the claims oversell the package to its detriment.”
Referee #2, second-round summary
The authors accepted both points in full. On interpretability: “we have removed claims of general ‘interpretability’ and now make the more defensible claims of (a) replicability across UKB, MGB, and AoU, and (b) clinical coherence with established biology in specific validated cases. We thank Referee #2 for pushing us toward this more accurate framing.” On leakage: “We accept this critique in full. Every instance of ‘leakage-free’ has been removed from the manuscript.” They added two limitations — the sixth, that a first-recorded diagnosis code is a proxy for onset and “documentation lag between true onset and EHR capture can be substantial — often several years for chronic conditions”; the seventh, on the scope of biological interpretation they can defend — and closed: “We believe the revised manuscript no longer oversells the method and now describes its scope honestly.” The published paper uses the adjective “interpretable” zero times. One “interpretability” survives, in the Methods, describing the choice of 21 signatures.
The release, dated 15 July at Dana-Farber and 17 August at Harvard, quotes Parmigiani: “We have painstakingly curated these signatures, which is a big differentiator [from other models]. In contrast to deep-learning approaches, which are typically ‘black boxes,’ our curated signatures capture the underlying biology in an interpretable way.” There are two ways to read “curated.” The charitable one is that it describes design effort — the choice of 21 signatures, the spectral-clustering initialisation, the 348-PheCode threshold, the 36 risk scores — and those are human decisions. The paper, though, describes the signatures as “inferred by ALADYNOULLI,” says their temporal relationships “emerge directly from the model without explicit encoding,” and says the approach “does not rely on disease-specific risk factors or manual feature engineering.” The release makes a broad interpretability statement. The journal response, six weeks earlier, removed general interpretability claims while retaining coherence claims for specific validated cases. The record does not identify who drafted the release or whether anyone compared the two.
“First known”: a sentence the paper never makes
The paper claims no priority. Its abstract says the model’s disease-level predictions “complement” Delphi-2M, a generative transformer from the German Cancer Research Center and EMBL published in Nature in September 2025 that predicts rates of more than 1,000 diseases from UK Biobank records and was evaluated, without retraining, in 1.9 million Danish registry records. The authors added the Delphi-2M comparison at Referee #3’s request, called its performance “comparable overall,” and — because Delphi-2M’s parameters were not available to them — compared against its published metrics rather than running it. Earlier multi-disease EHR predictors go back at least to BEHRT in 2020. What Delphi-2M and its predecessors do not do is put genetics inside the model: Delphi-2M takes diagnosis codes, sex, BMI, smoking and alcohol, and touches polygenic scores only in a post-hoc analysis. So the release’s qualifier — “based solely on patient health records and genetic profiles” — does narrow the claim to something the paper’s own comparator does not cover. Whether any earlier model jointly fitted records and polygenic scores across hundreds of diseases is contestable; a 2025 preprint the supplement itself cites, on adding polygenic scores to an EHR foundation model, sits close. We treat “first known” as unsupported by the paper and unverified; we do not call it false.
On “AI”: the model is fitted by gradient optimisation in PyTorch, so “machine learning” is defensible umbrella language, and the paper itself sets the method against “black-box machine learning methods,” which places it inside the category. We put little weight on the word.
What happens to the one-year AUC when the last year of records is withheld
Referee #2’s leakage point has a number attached. ICD codes are often entered well after a disease begins, so a model that reads records up to the moment of prediction may be reading the early paperwork of the disease it is predicting. The authors ran three washout analyses. Excluding outcome events occurring within 1, 3 or 6 months of enrolment lowers the mean one-year AUC by 0.0112 — small. Excluding the first year of outcomes lowers the ten-year ASCVD AUC from 0.733 to 0.723 — small. The third analysis withholds the final one or two years of records before each annual prediction point and asks for a one-year prediction. For ASCVD, the AUC goes from 0.85–0.90 to 0.69–0.75; for diabetes from 0.76–0.82 to 0.57–0.70; for atrial fibrillation from 0.69–0.87 to 0.49–0.76. Those values are printed in Supplementary Figure 19c, described in the prose only as showing “substantial residual predictive power.” The authors’ response to reviewers tabulates three more diseases that did not make the published figure, including breast cancer, where the same exercise takes 0.975 and 0.963 at the first two timepoints down to 0.695 and 0.701 — figures that appear only in that response and are not reconciled with the lower enrolment and median values in Table S14 (0.782 and 0.867).
| Disease · records withheld before the prediction point | +1 yr | +2 | +3 | +4 | +5 | +6 | +7 | +8 | +9 |
|---|---|---|---|---|---|---|---|---|---|
| ASCVD · none | 0.848 | 0.894 | 0.878 | 0.885 | 0.902 | 0.876 | 0.87 | 0.898 | 0.872 |
| ASCVD · 1-year washout | 0.735 | 0.754 | 0.714 | 0.728 | 0.747 | 0.709 | 0.686 | 0.713 | 0.716 |
| ASCVD · 2-year washout | – | 0.747 | 0.701 | 0.71 | 0.735 | 0.701 | 0.668 | 0.715 | 0.707 |
| Diabetes · none | 0.782 | 0.816 | 0.81 | 0.792 | 0.758 | 0.783 | 0.799 | 0.808 | 0.767 |
| Diabetes · 1-year washout | 0.655 | 0.699 | 0.677 | 0.639 | 0.618 | 0.651 | 0.599 | 0.651 | 0.568 |
| Diabetes · 2-year washout | – | 0.7 | 0.674 | 0.642 | 0.615 | 0.647 | 0.608 | 0.652 | 0.572 |
| Heart failure · none | 0.683 | 0.802 | 0.814 | 0.878 | 0.807 | 0.85 | 0.798 | 0.837 | 0.839 |
| Heart failure · 1-year washout | 0.636 | 0.714 | 0.74 | 0.747 | 0.668 | 0.748 | 0.675 | 0.688 | 0.703 |
| Heart failure · 2-year washout | – | 0.715 | 0.734 | 0.752 | 0.669 | 0.748 | 0.664 | 0.678 | 0.697 |
| Atrial fibrillation · none | 0.784 | 0.805 | 0.866 | 0.832 | 0.844 | 0.841 | 0.721 | 0.766 | 0.691 |
| Atrial fibrillation · 1-year washout | 0.673 | 0.644 | 0.757 | 0.662 | 0.649 | 0.644 | 0.527 | 0.572 | 0.489 |
| Atrial fibrillation · 2-year washout | – | 0.641 | 0.749 | 0.669 | 0.66 | 0.653 | 0.525 | 0.571 | 0.494 |
- This is context. It is not a verdict. Withholding the final year of diagnoses lowers every one-year AUC — ASCVD from 0.85–0.90 to 0.69–0.75 — but that cannot apportion the change between legitimate recent clinical information and leakage through mis-dated diagnosis codes, and a one-year washout also changes the clinical task being scored. The authors' own limitation 6 says documentation lag means some nominally incident events are prevalent disease.
- Cut the other way too: the ten-year ASCVD AUC barely moves when the first year of outcomes is excluded (0.723 versus 0.733), and excluding outcome events within 1, 3 or 6 months of enrolment lowers the mean one-year AUC by 0.0112 (Methods).
- The paper's prose reports these multi-year results only as “substantial residual predictive power”; the numbers sit in the supplementary figure. The authors' response to reviewers tabulates three more diseases — breast cancer (0.975 and 0.963 at years 1 and 2 without washout; 0.695 and 0.701 with a one-year washout), all cancers and chronic kidney disease — that do not appear in the published figure; the breast-cancer values are not reconciled with the lower enrolment and median values in Table S14 (0.782; 0.867).
- A direct comparison with the ten-year clinical scores (PCE 0.683, QRISK3 0.702) would be improper — different horizon, different task. The panel shows only how the one-year AUC changes when recent records are withheld.
- Atrial fibrillation at year +9 with a one-year washout is 0.489 — below chance — on a single 10,000-person batch; the paper gives no cell-specific event count, so the tail cells should not be over-read.
This cuts both ways. A model that uses the most recent year of a person’s record is using information a clinician would also use; the short-washout and ten-year results barely move. The washout shows how the one-year AUC changes when recent records are removed; it cannot apportion that change between valid recent information and temporal leakage through mis-dated codes, and because a one-year washout defines a different clinical task it cannot support a direct comparison with the ten-year clinical scores. What it does show is that the striking one-year figures — the 0.88 for heart disease, the 0.97 for breast cancer — change substantially when the final year of records is withheld. The authors’ sixth limitation acknowledges the diagnosis-date uncertainty; the figure quantifies performance under washout; the release mentions neither.
In fairness
- The paper is open, and so is its review. Open access, a 48-page supplement with confidence intervals for the principal one-year-at-enrolment and ten-year AUCs (the median one-year and Cox baseline values have none), and an 88-page peer-review file in which the authors accept the two critiques quoted here in writing. This scrutiny is possible because the authors and the journal published the material.
- The discovery package is substantial. Signatures that replicate across three health systems, a signature-level GWAS, rare-variant and carrier validations, and inverse-probability weighting for the UK Biobank’s healthy-volunteer bias. Referee #3: “a really strong paper.” Referee #2, in the same report that objected to the claims: “an interesting method that is well-execute on a number of levels” [sic].
- The ten-year heart-disease result is real. Same participants, same events, the scores at the horizon they were built for, and the model ahead by 0.03–0.07 with non-overlapping intervals.
- The evaluation design is prospective in form. Forty held-out batches; population parameters fixed; only each person’s own trajectory fitted on the held-out batch; predictions from data available at the prediction point. One qualification the paper states and does not quantify: the spectral-clustering initialisation of the signature–disease starting values was run once on the entire dataset and reused in every batch.
- The release carries a real caveat — “only been tested on past patient data” — restricts the Gail claim to one year, and matches the paper’s funding list line for line. Mass General Brigham’s own Q&A on the paper makes no “first” claim and describes “21 reproducible, ‘latent’ disease signatures.”
Traceability, as of 26 August 2026
The paper’s code-availability statement says all code is on GitHub under the “MGB Source Available License” and that “a versioned snapshot tagged v1.0 at the time of acceptance is permanently archived on Zenodo.” On the dates we checked (19, 23 and 26 August 2026):
- The public repository,
surbut/aladynoulli2, holds one commit, dated 17 June 2026 — nine days after acceptance — and no tags or releases. - The cited Zenodo record (10.5281/zenodo.20802505) is publicly listed but its files are restricted; a second Zenodo record cited in the Reporting Summary (10.5281/zenodo.20187989) is marked deleted in a public tombstone dated 21 July 2026, which gives “Personal data issue” as the reason and points to a replacement DOI (10.5281/zenodo.20802504).
- The licence permits research use only and bars distribution of modified versions; the repository README displays an MIT badge.
- No code computing PCE, PREVENT, QRISK3 or Gail is in the repository; trained model parameters and checkpoints are not released (the authors told reviewers data-access restrictions prevent it); the documented batch-training entry point failed on import when we ran it on 19 August 2026, because a module it references is absent from the repository.
- The web application the authors described to reviewers as “available for clinical use” was, when we loaded it on 19 and 23 August 2026, a demonstration labelled “SYNTHETIC / DEMONSTRATION DATA ONLY — NOT REAL UK BIOBANK DATA”; on 26 August the site did not respond.
This can change quickly; we will update this box if it does.
What would move this rating
- Prediction results from Mass General Brigham and All of Us. The Reporting Summary says model performance was assessed separately in those biobanks; the paper says no prediction tasks were performed there. Published external prediction results — per-disease discrimination, calibration and operating thresholds — would be the single strongest upgrade.
- A genetics-free ablation. If the model predicts nearly as well from diagnosis codes and age alone, the “existing patient data” framing holds; if the 36 risk scores carry the load, it does not.
- A threshold, a sensitivity and a positive predictive value for the colorectal scenario — or any disease — so that “high accuracy” means something a clinician can act on.
- Per-disease calibration, and a head-to-head Delphi-2M run now that its weights are distributed through the UK Biobank.
- A narrower “first.” The claim that survives the record is something like “the first to fit diagnosis histories and polygenic scores jointly in one generative model across hundreds of diseases.” Said that way, we would have no quarrel with it.
The bottom line
Believe this: a Bayesian model that reads a person’s diagnosis history and polygenic risk scores can produce one-year risk rankings that are strong for some diseases and, for heart disease at ten years, better than the clinical scores in use — on 400,000 UK Biobank volunteers, in a careful, open, peer-reviewed study whose authors accepted the interpretability and diagnosis-date critiques quoted above and added limitations and sensitivity analyses in response.
Don’t believe this: that a new AI tool has been shown to predict more than 300 diseases from existing patient data. What has been shown is discrimination for 28 diseases in one biobank, in genotyped people, with ten-year figures that are modest for most conditions — eleven at or below 0.60, depression’s interval wholly below 0.5 and bipolar disorder’s point estimate below it; a colorectal-cancer “high accuracy” that is an AUC with no referral analysis behind it; and a “first” the paper does not claim. The adjective “interpretable” is absent from the final paper and appears in the release. The paper’s own limitations describe the tool more accurately than its press did.
Disclosures & provenance
- Published
- 26 Aug 2026
- Author
- The Ground Truth editor. Editorial standard →
- Funding
- Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
- Rating
- Overstated (2/5), rubric v1.0, rated 26 Aug 2026. Full rating card →
- Corrections
- None to date. Corrections log → · Challenge this analysis