Read before citing
  • The underlying study is real, large, prospective and open access, and the gynaecological result is among its strongest. Nothing here says the test does not work. This is about the distance between what was measured and what was said.
  • Nothing was acted on. Clinicians were blinded to every result and, in the paper’s words, “the patient’s onward journey was unaffected.” The study therefore provides no evidence that any procedure was avoided because of a PinPoint result — it did not evaluate procedure avoidance at all.
  • The two numbers behind “99%” are sensitivity, 99.1% and negative predictive value, 99.8%. They answer different questions about different groups of women and cannot be added together into a symmetric “99% accurate at detecting and ruling out.”
  • The “one in five women ruled out” is close to the setting the researchers chose, not a discovery. The operating point is defined as a threshold that clears 20% of the women who do not have cancer.
  • The “18,000 women a year” figure, the transvaginal-ultrasound comparison and the AUC of 0.832 versus 0.792 in 578 women appear nowhere in the paper or its supplement. They appear on one NHS Innovation Accelerator page and we could not trace them to any published study, dataset or statistical report.
  • Every figure here was re-derived from the paper’s own tables through a second, independent pass — the paper retrieved again from a separate archive and the arithmetic recomputed in code, run against this draft rather than with it. It required corrections, including one error of fact. They are described at the end.

What was claimed

On 8 July 2026 the Guardian reported that thousands of women could be spared a painful examination by a new NHS AI blood test. The load-bearing sentence read: “The results showed that the test had a 99% accuracy rate in both detecting the gynaecological cancers found among the 3,313 women and also ruling out its presence – a higher success rate than conventional testing.” The article added that the test “could save one in five of those women – 18,000 a year – from needing to undergo a diagnostic procedure called a transvaginal ultrasound scan.”

The test is PinPoint, made by PinPoint Data Science in Leeds. It is not an imaging model or a new assay. It is a machine-learning model that takes 31 blood analytes measurable on standard NHS laboratory platforms, plus age and sex, and returns one of three risk bands. No new analyser is needed. The paper does not establish whether the full panel is already ordered for every referral or whether implementation needs a dedicated sample. That is a genuinely attractive proposition, and part of why the result deserves careful reading rather than dismissal.

The study behind it was published in Mayo Clinic Proceedings: Digital Health as “Real-World Validation of PinPoint Blood Tests in the NHS”. It is open access under CC BY 4.0. It reports nine separate tests, one per urgent-referral pathway, across 16,481 patients enrolled at five NHS trusts and 170 GP surgeries between December 2020 and July 2025. The gynaecological result is one of nine, and the paper judges five of the nine to show potential clinical utility.

Not one of the nine news reports, company statements or NHS pages we examined mentions the fact that shapes everything else: the clinicians treating these patients never saw the results.

One cohort, five numbers

The gynaecological analysis covers 2,953 patients — 2,952 women and one man, who did not have cancer, so every cancer below is a woman’s. Of those patients, 114 were diagnosed with cancer. The paper reports three operating points, and the supplement gives the raw counts behind each. Here they are.

2,953 patients · gynaecological referral pathway · PinPoint v1.2

One cohort. One model. Five different “accuracy” numbers.

Had cancerDid not have cancer
Flaggedelevated / high risk
113
correctly flagged
2,272
false alarm
Not flaggedlowest risk
1
test false negative
567
correctly cleared

Both of the numbers that became “99%” are in there, and they are not the same number. 99.1% is sensitivity: of the 114 women who had cancer, 113 were flagged. 99.8% is negative predictive value: of the 568 women placed in the lowest-risk band, 567 did not have cancer. The first is a fact about women with cancer. The second is a fact about women the test cleared. Fusing them into “99% accurate at detecting and ruling out” produces a sentence that sounds symmetric and is not.

It also quietly disposes of the third number. Specificity was 20.0%. Four out of every five women without cancer were still flagged as elevated or high risk. That is not a defect — it is what a rule-out tool set to this threshold does, and we return to why below — but it is the opposite of what a reader takes from “99% accurate at ruling it out.”

And there is a fifth number, which no one has published because the paper does not calculate it. “Diagnostic accuracy” is used loosely in medicine as an umbrella for the whole family — sensitivity, specificity, predictive values, AUC. But read the way a classifier is normally scored, as overall classification accuracy, the proportion of all classifications that were correct, the operating point that produced the 99.1% and the 99.8% got 680 of 2,953 right. That is 23.0%.

We are not offering 23% as the real score, and it would be unfair to. It is a poor way to judge a rule-out policy — which is exactly why the paper reports sensitivity, specificity and predictive values instead, and why the authors were right to. We calculate it only to make one point, which we think is the useful one for a reader: “accurate” does not pick out a measure, and this test does not have a single number behind it. On the classification reading it scores 23%, where flagging nobody would score 96.1%. So the problem with the headline was never that it chose the wrong X. It is that a reader had no way to know which measure either 99% referred to — and those two 99s answer questions that matter very differently to a woman waiting for a scan.

One cancer

The most important number in this study is not a percentage. Of the 114 diagnosed gynaecological cancers, one received a lowest-risk PinPoint score — one test false negative. That is a good result. It is also a fragile one: the entire rule-out case rests on a single observed false negative, which is why the confidence interval on that sensitivity runs from 95% to 100%.

This was not a clinical miss. No clinician saw the score, so that patient went through the unchanged referral pathway and was diagnosed by it — which is how the study knew to count her as a cancer at all. What the paper does not report is the site or stage of that cancer, which investigation found it, or what would have happened had the score been allowed to change her care. It reports no stage data anywhere.

Nor is the cancer mix what “gynaecological cancers” suggests. From the supplement’s ICD-10 table:

Cancers found on the gynaecological pathway, by ICD-10 code
SiteICD-10CasesShare
Endometrial (corpus uteri)C546960.5%
VulvaC5197.9%
OvaryC5697.9%
Retroperitoneum / peritoneumC4887.0%
VaginaC5254.4%
CervixC5343.5%
Uterus, unspecifiedC5521.8%
Non-gynaecological, found on this pathwayvarious65.3%
Other21.8%

The evidence is dominated by endometrial cancer, as the authors themselves say: the test was “validated primarily against endometrial cancer, 61% of cohort cancers.” With nine ovarian cancers and four cervical, the study cannot establish site-specific rule-out performance for either, and does not claim to. The plural in the headlines does work the data does not support. PinPoint’s own chief medical officer, quoted in the coverage, narrowed it in the other direction — to “more than 99% of endometrial cancers” — which the paper also does not report, because it publishes no endometrial-only sensitivity.

“One in five” was the setting, not the finding

This is the part that most changes how the story reads. The paper does not discover that one in five women can be safely cleared. It defines the operating point that way. From the Methods, the first of three illustrative use cases is:

20% Rule-Out: threshold set so that 20% of disease-negative patients with the lowest test scores are identified as test-negative.

The threshold is chosen to clear a fifth of the women without cancer. That is why specificity reads 0.20 in every single pathway in the paper’s Table 2 — breast, lung, skin, urological, all of them. It is an instruction, not a result.

What follows from that instruction is a finding, and we should be precise about it: 568 of 2,953 patients — 19.2% — landed in the lowest-risk band, and only one of them had cancer. That is the real, earned result. But “researchers say this could allow one in five symptomatic women to be safely ruled out,” as the NHS Innovation Accelerator put it, presents the dial setting as though it were the discovery.

Move the dial and everything moves with it. Set it to catch 90% of cancers instead and the same model on the same women misses twelve rather than one. Set it to prioritise the highest-risk tenth — the use the paper’s own abstract leads with — and it finds 41% of cancers. Same test, same blood, same day.

Where 360 women went

Two different cohort sizes have been circulating: 3,313 and 2,953. Both are real, and the difference matters.

Figure 1 of the paper gives the flow. 16,855 patients were reported recruited; 16,481 were successfully tested and enrolled, of whom 3,313 were on the gynaecological pathway. Then 29 failed the age criterion and 3,197 were excluded — most commonly because their data could not be extracted from NHS systems (1,470) or their referral was never completed (682). That leaves 13,255 in the primary analysis, of whom 2,953 were gynaecological.

So 3,313 is the enrolled count and 2,953 is the analysed count. Every performance figure in this article, and every figure quoted in every news report, comes from the 2,953. The Guardian attached the 99% to “the 3,313 women”; the NHS Innovation Accelerator said the test was “assessed across” 16,481 referrals including 3,313 women. Both used the larger number as the denominator for statistics computed on the smaller one. 360 gynaecological patients — 10.9% of those enrolled — appear in no performance figure at all, and the paper does not publish a breakdown of who they were.

There is a deeper question underneath all of this, and it is the one we would put first to a methodologist. Cancer status here comes from the outcome of the routine referral, generally completed within three months; a patient with no recorded cancer diagnosis was assigned to the non-cancer group. There was no uniform pathology reference standard and no long-term follow-up. Blinding protects against the score changing the work-up, but it does nothing about a cancer that was present and simply not recorded in the window. Because the whole rule-out case rests on a single false negative, even a handful of occult or late-recorded cancers among the 568 cleared patients would move it materially. Nobody has published that audit.

The exclusions are not obviously biased, and the authors ran a worst-case sensitivity analysis suggesting the effect on AUC is at most 0.08 in the worst-affected pathway and under 0.06 elsewhere. But that analysis addresses discrimination, not the miss count — and the miss count is one. The largest exclusion category, patients whose data could not be extracted, has no recorded cancer outcome at all.

The Independent and Medscape UK both used 2,953. They were right.

The road to 18,000

The paper contains no cost model, no estimate of procedures avoided, and no figure of 18,000. Its single downstream sentence points somewhere else entirely: the gynaecological test “could reduce demand for hysteroscopy, of which ∼100,000 are performed annually in England with a 5%-10% cancer yield.” Hysteroscopy, not ultrasound; “could,” not will; and, as it turns out, an undercount — NHS England’s 2024/25 cost collection records about 176,000 diagnostic hysteroscopies. That sentence carries no citation.

The 18,000 is press-side arithmetic: roughly 90,000 annual referrals for post-menopausal bleeding, times one in five. We could not verify the 90,000 against any NHS collection, because none exists at that granularity — the cancer-waiting-times data has a single “suspected gynaecological cancer” category, and it recorded 319,798 urgent referrals in 2025–26, about three and a half times the figure in the claim.

Every link between the finding and the headline

Six things must hold for 18,000 women to skip a scan. One was measured.

  1. Proxy figure Women referred each year on the suspected-womb-cancer pathway in England

    The claim puts this at about 90,000. PinPoint’s press release does footnote it — but not to any count of post-menopausal-bleeding referrals: to hysteroscopy volumes (71,000 in England in 2019–20, scaled for population growth to about 87,000 UK-wide in 2026) and to NHS Scotland guidance. Hysteroscopies are procedures performed, not referrals received. NHS England’s cancer-waiting-times collection has no womb-specific category to check it against; its one gynaecological category recorded 319,798 urgent referrals in 2025–26.

  2. Set by the threshold …of whom about one in five score in the lowest-risk band

    Not a discovery: the operating point is defined as clearing 20% of the patients without cancer. In the study 568 of 2,953 patients — 19.2% — fell in that band.

  3. Assumed …and are eligible, with a usable blood result

    In the evaluation, 360 of 3,313 enrolled gynaecological patients — 10.9% — never reached the analysis.

  4. Never tested …and their clinician accepts the low-risk score and cancels the scan

    Clinicians in this study were blinded to the result, so this was never observed once. The only published PinPoint economic model — for urological cancer — simply assumes low-risk patients are not referred, with no adherence parameter at all.

  5. Never tested …and has no other reason to be scanned

    Persistent bleeding, age, tamoxifen, obesity or Lynch syndrome all still point to imaging. And the scan being replaced does more than measure the endometrium: it images the ovaries and the pouch of Douglas. Seventeen of the 114 cancers here were ovarian (nine, C56) or coded to retroperitoneum and peritoneum (eight, C48).

  6. Never tested …and no cancer is diagnosed late as a result

    The study followed referrals, not the women who would have been sent home. At the miss rate observed inside its own rule-out group — one cancer in 568 — clearing 18,000 women a year would leave roughly 32 cancers in the lowest-risk band annually (95% CI 6 to 178; the interval is that wide because it rests on one observed miss).

Procedures avoided on these assumptions

— the setting the published figure assumes.

The published number is what you get when every unmeasured link passes at 100%. That is a ceiling, not an estimate. Nothing here says the ceiling is unreachable — only that the study measured the first step and assumed the rest.

There is a further wrinkle in the substitution. NICE’s NG12 does not, in fact, require a transvaginal ultrasound; it is a referral guideline. The ultrasound-first rule comes from British Gynaecological Cancer Society guidance, which uses a 4mm endometrial-thickness cut-off to decide who needs a biopsy or hysteroscopy. So the scan PinPoint would displace is the cheap triage step that decides whether the expensive, invasive ones happen — and it is also the step that looks at the ovaries.

The comparison that is not in the paper

One striking head-to-head circulated with the story: in a subgroup of 578 women, PinPoint achieved an AUC of 0.832 against 0.792 for endometrial thickness measured by ultrasound. It is the only direct comparison with current care anywhere in the story. The Guardian printed the vaguer “a higher success rate than conventional testing” without those figures.

It is not in the paper. There is no 578-woman subgroup, no endometrial-thickness comparison and no AUC of 0.832 or 0.792 in the article or its supplement. We searched the full text, every supplementary table cell, and the figures — then repeated the search on a separately retrieved copy of the paper, and found the same. A full-text search of Europe PMC for those two AUC values alongside “endometrial” returns nothing. The Accelerator page is itself a published source; what we could not find is a published study, dataset or statistical report behind it.

The figures appear in two places, neither of them a study. An NHS Innovation Accelerator news item of 8 July introduces them as “a smaller subset of 578 women” inside continuous prose about the same NHS evaluation, and attributes them to no source at all: no journal, no DOI, no authors, its only outbound link the Guardian. PinPoint is one of the Accelerator’s own portfolio innovations. PinPoint’s own press release of the same day carries the identical comparison — and does attribute it, to “PinPoint sub-analysis, data on file.” That is the company telling you, in its own footnote, that the number behind the story is unpublished.

The comparison the paper does make points the same way but far more modestly, and the authors flag its limits themselves: CA125 alone scored an AUC of 0.71 (0.66–0.77) on this pathway against PinPoint’s 0.81 (0.77–0.85). In their words: “While this is an imperfect comparison, it suggests that there may be a role here for the gynecological test.” It is worth knowing why it is imperfect: CA125 is not an outside rival but one of the model’s own 33 inputs, so this is a multivariable model measured against its own best single feature. That is a fair thing to report and a weak thing to build a superiority claim on.

Two models

There is a wrinkle the coverage missed entirely, and the authors did not hide it — they report it in the open, explain its cause, and test for its consequences.

Two versions of the algorithm were scored on the same 2,953 patients and the same blood draws. Version 1.1 was the pre-existing model. Version 1.2 was developed in parallel with the evaluation, as the protocol allowed, after interim analyses revealed that missing-not-at-random tumour-marker values in the original training data were distorting predictions. It was retrained on the same retrospective Leeds dataset with imputation and calibration corrections, and locked before any final analysis.

Two details matter here and both favour the authors. The interim analyses that triggered the redevelopment concerned the lower-gastrointestinal, breast, skin and urological pathways — gynaecology was not among them. And those four pathways closed to recruitment in February or July 2023, while gynaecological enrolment ran until July 2025. So the decision to build v1.2, and the choice of what to fix, was driven by looks at other pathways. That materially reduces the risk of leakage into this result — though the paper does not give the exact date v1.2 was locked, or how much of the gynaecological cohort had accrued by then.

Gynaecological pathwayAUC (95% CI)Sensitivity at rule-outTest false negatives
v1.1 — pre-existing model0.69 (0.63–0.74)0.956 of 114
v1.2 — produced every headline figure0.81 (0.77–0.85)0.991 of 114

For context, the gynaecological model scored an AUC of 0.8124 in the retrospective validation reported in the 2022 development paper — a figure that appears only in that paper’s supplement. So the pre-existing model, meeting prospectively collected patients for the first time, fell about 0.12 below its own retrospective result; the corrected model returned to it.

The authors’ defence is reasonable and they make it explicitly: v1.2 was locked before final analyses, and the training data was the old retrospective set rather than the evaluation cohort. They also run a leakage check — though it is worth being precise about its scope, since they are: the no-leakage sentence covers the breast, skin, urological and lower-gastrointestinal pathways, the four that prompted the redevelopment, and gynaecology is not among them. A gynaecological temporal result does exist separately (0.82 then 0.81, p=0.69), but it found no evidence of a difference rather than demonstrating equivalence, and the authors do not offer it as a leakage defence for this pathway.

Two things follow that no press account conveyed. Every headline number describes v1.2. And the version question is answerable only from a document no coverage cited: PinPoint publishes UK Declarations of Conformity on its own site, and the gynaecological one (REC-777, signed by the chief executive and chief scientist on 13 May 2026) declares that “PinPoint_Gynae v1.2 is in conformity with the essential requirements and provisions of UK SI 2002 No. 618 (Medical Devices Regulations); we therefore affix a UKCA mark to this product.” So the marked product is the corrected model — and the same document records the notified body and certificate number as “N/A (Self-Certified Device).” The MHRA register, where a reader would think to look, still carries no version field at all.

Who has checked it

Eight of the eighteen authors are employed by PinPoint Data Science and hold shares or options in it. The University of Leeds and Leeds Teaching Hospitals Trust hold a royalty agreement with the company, and one author is a named inventor under it. The company funded seven authors’ contributions. All of this is disclosed in the paper, plainly. We did not uncover it; we are repeating it, because it changes what independent verification is still needed rather than what the study found.

What is harder to see from any single paper is the shape of the whole evidence base. There are four published evaluations of PinPoint — the 2022 development study, a breast-clinic simulation, a urological economic model and this one. Every one has at least one author with a direct financial interest, and one author appears on all four. The breast simulation was funded by an Innovate UK grant led by PinPoint itself; the urological economic model was funded by the same SBRI Healthcare award as this study, and took its accuracy inputs from the company’s own earlier paper. We could find no independent replication of PinPoint by a team with no financial relationship to it. We could find no NICE assessment either.

“Fully regulated” is also doing more work than it looks. The phrase is PinPoint’s own, from a commercial profile in Open Access Government credited to the company’s chief communications officer. PinPoint is registered with the MHRA as an In Vitro Diagnostic Device of sub-type “IVD General” — the lowest UK category, which manufacturers may self-certify without any approved body assessing performance. The register records no performance studies and has no field for the algorithm version; as of 27 July 2026 it lists nine PinPoint products, none of them carrying a version number.

How the number changed hands

Nine organisations restated this result within two weeks. Several got it right, and the difference between those that did and those that did not is the most useful thing in this piece for anyone reading the next such story.

Every published restatement, in order · 7–21 July 2026

What happens to a number in transit, from the table to the headline.

Every restatement is shown in full below. Tap a change to see which ones carried it — or to scan the pattern.

    Three outlets got it right, and they did the same thing: they kept the paper’s own words. “Correctly identified 99.1% of cancers as elevated or high risk” is a sentence you cannot compress into “99% accurate” without losing the meaning.

    The company supplied the correct technical wording to reporters, and several outlets kept it. But PinPoint’s chief medical officer — a co-author of the paper, an employee and an equity holder — also used “accuracy”, in an endometrial-cancer formulation. The wording was not confined to journalists.

    Rating the parts separately

    “99% accurate” travelled with four other claims, and they do not stand or fall together. Our rating at the top of this page covers the accuracy claim itself — the one the headline asks about. The rest are rated here, separately, because averaging them would tell you less than listing them:

    The part of the claimBandWhy
    “99.1% of cancers flagged as elevated or high risk; negative predictive value 99.8%”Holds up with conditionsNumerically exact. The conditions: model v1.2, 2,953 analysed rather than 3,313 enrolled, one test false negative, achieved at a specificity of 20.0% and a PPV of 4.7%, and an NPV that moves with prevalence.
    “AUC 0.81, better than CA125’s 0.71”Holds up with conditionsBoth figures are the paper’s. CA125 is one of the model’s own inputs, and the authors call the comparison imperfect. Say so.
    “One in five women could be safely ruled out”Needs contextThe 20% is the threshold’s definition; the observed share is 19.2%. “Safely” is asserted, never demonstrated — nobody states what miss rate would count as acceptable.
    “A trial”OverstatedAn approved NHS service evaluation with verbal consent and no trial registration. The word “evaluation” appears nowhere in the Guardian’s account.
    “99% accurate at both detecting and ruling out” — the claim rated at the top of this pageOverstatedBoth 99s are real and exact, but they measure different groups of women. Fused into one symmetric figure they conceal a specificity of 20.0%. The quantity the word names is 23.0%.
    “A higher success rate than conventional testing”MisleadingNo such comparator is in the paper, and it is unsourced in every outlet that printed it.
    “18,000 a year spared a transvaginal ultrasound”MisleadingNot derivable from the study, names a procedure the paper never mentions, and assumes an action the study never took.
    “20% of women referred turn out not to have the disease”MisleadingThe true figure is 96.1%, and the same article contradicts it two sentences later.
    The test itselfnot rated hereOn the paper alone the gynaecological result would plausibly sit at the top of this table. We rate claims, not products.

    One consequence of the prevalence point deserves spelling out, because it cuts against the claim in its own terms. The 99.8% negative predictive value was measured where 3.9% of women had cancer. The claim applies it to post-menopausal-bleeding referrals, which the coverage itself says are about 10% cancer. Run the study’s own sensitivity and specificity at that higher rate and the negative predictive value falls to 99.5% — and the number of cancers sitting inside the reassured group rises from roughly 30 a year to roughly 80.

    What the study actually earns

    Strip away the sentence and there is a real result underneath, and it deserves stating as plainly as the criticisms.

    This is the first large, prospective, multi-site evaluation of this technology in NHS England: five years, five trusts, 170 GP practices, with staff asked to enrol every eligible patient, and cancer status taken from the outcome of the routine referral rather than assigned by the researchers. The blinding that makes the “women spared a scan” framing impossible is also the study’s central methodological strength — with no clinician able to see a score, no feedback loop could flatter it.

    The gynaecological model achieved an AUC of 0.81 (0.77–0.85) on a pathway spanning at least seven cancer sites, using only analytes measurable on standard NHS platforms. It was trained on data from a Siemens laboratory at Leeds and evaluated on samples run at a Roche/Sysmex laboratory at Mid Yorkshire. Analyte values are analyser-dependent, so holding discrimination across a complete platform change is a genuine external analytical validation, and one most published models never attempt. Performance was stable across the five-year window (0.82 then 0.81, p=0.69). Calibration-in-the-large was favourable: the paper names four pathways whose observed-to-expected interval excluded 1 and gynaecology is not among them. That does not by itself establish calibration across the whole score range. Decision-curve analysis showed net benefit equal to or better than the treat-all and treat-none comparators at every threshold displayed — modelled net benefit, not observed patient impact. And 113 of 114 cancers landed above the rule-out line.

    A 20% specificity is not the scandal it looks like either. On a pathway where 96% of referred women do not have cancer, a tool whose job is to find the safest fifth to reassure is not trying to be right about everyone. The right question is not “why is specificity so low” but “is the lowest-risk band safe enough to act on, and does acting on it help.” The first of those the study makes a decent start on. The second it does not touch.

    A headline that kept the meaning was available, and it is barely longer. Something like: Blood test flagged 113 of 114 gynaecological cancers in NHS evaluation — but clinicians never saw the results, and four in five women without cancer were flagged too. That is the finding, and it would have survived contact with the paper.

    What the evidence does not settle yet

    These are not rhetorical gaps. They are specific, and most are the kind only a different study can close — which is the honest position for a technology at this stage.

    What we would still want to know

    These are the questions the evidence raises that only the authors, the company or a further study can settle. We publish them because they are the questions a reader should want answered before this test changes anyone’s care — and because if any of them are answered, this article should change. Answers, from any source, will be published in full and logged under our corrections policy.

    1. What was the site and stage of the single gynaecological cancer classified as lowest risk?
    2. PinPoint attributes the 578-woman comparison against endometrial thickness, with AUCs of 0.832 and 0.792, to an unpublished sub-analysis held as “data on file.” Will it be published, with its methods and its cohort definition?
    3. Which model version is running at the trusts now adopting the test? The UKCA-marked product is v1.2, per the company’s own declaration of conformity; what is deployed is not documented anywhere we could find.
    4. What assumptions generate the estimate of 18,000 avoided transvaginal ultrasounds, and is there a direct count of post-menopausal-bleeding referrals to put behind the 90,000 in place of scaled hysteroscopy volumes?
    5. What is the source for “∼100,000 hysteroscopies annually in England” in the Implications section?
    6. Is there a post-menopausal-bleeding subgroup analysis, given that this is the population the claim addresses?
    7. Will the per-site performance breakdown be published?
    8. What is the AUC of an age-only, or age-and-sex-only, model on the gynaecological pathway?
    9. How many of the 568 patients in the lowest-risk band had definitive pathology, as opposed to a referral closed with no cancer code recorded within the window?
    10. On what date was v1.2 locked, and how many gynaecological patients had been recruited by then?
    11. In Supplemental Table 5, the v1.1 sensitivity interval is given as 0.88–0.98; a standard Wilson interval on 108/114 gives 0.89–0.98. Which method was used?

    The bottom line

    The test may well be good. On the evidence published so far it discriminates better than the marker it would sit alongside, it runs on analytes NHS laboratories already measure, and it put 113 of 114 cancers above its rule-out line while clearing nearly a fifth of the pathway. Those are real achievements and they were earned prospectively, which is more than most health-AI claims can say.

    The sentence, though, is not the test. “99% accurate at detecting and ruling out” fuses two different measurements of two different groups of women, discards a specificity of 20%, and lands on a word whose actual meaning here is 23%. “One in five women ruled out” reports a dial setting as a discovery. “18,000 women spared a scan” is an unmeasured projection built on an unsourced denominator, describing a procedure the paper never mentions, in a study that never evaluated procedure avoidance, because no clinician was ever shown a result.

    If you take one thing: ask which 99%. The answer changes who the number is about, and it is the difference between a test that finds nearly every cancer and a test that clears nearly nobody.

    Method

    Every figure here comes from the paper and its supplementary appendix rather than from coverage of them. The paper is open access and free to read, and the supplement publishes the cross-tabulation behind every operating point — which is what makes the confusion matrix in this piece reconstructable at all. Anyone can check it.

    Counts are transcribed from Table 1, Tables 2–4, Figure 1, and Supplemental Tables 3, 5, 8, 11–15 and 17–18. Figures we derived ourselves rather than took from the paper — overall classification accuracy, the flag-nobody baseline, and the scaled miss estimates — are shown as fractions wherever they appear, so the arithmetic is visible. All of it was re-derived a second time from an independently retrieved copy of the paper, with the calculations recomputed in code.

    National denominators come from NHS England’s Cancer Waiting Times collection (319,798 gynaecological urgent referrals, 2025–26) and the National Cost Collection 2024/25 (176,244 diagnostic hysteroscopies; about 582,000 transvaginal ultrasounds). Both are floors rather than ceilings.

    Primary source: Neal M, Dean M, Duffy S, et al. “Real-World Validation of PinPoint Blood Tests in the NHS: Multivariable Machine Learning to Predict Cancer Risk in Primary Care Urgent Referrals.” Mayo Clinic Proceedings: Digital Health 2026;4(3):100382. doi:10.1016/j.mcpdig.2026.100382. Open access, CC BY 4.0. Development paper: Savage R, Messenger M, Neal RD, et al. BMJ Open 2022;12:e053590.

    Confused by sensitivity, negative predictive value and AUC? Read our field guide: How to read a “99% accurate” diagnostic-AI claim.

    Disclosures & provenance

    Published
    25 Jul 2026 · last updated 27 Jul 2026
    Author
    The Ground Truth editor. Editorial standard →
    Funding
    Self-funded. Ground Truth takes no money from, and has no affiliation with, any organization examined here. Independence policy →
    Rating
    Overstated (2/5), rubric v1.0, rated 25 Jul 2026. Full rating card →
    Corrections
    Corrected 27 Jul 2026 — We wrote that no public record stated which model version is the UKCA-marked product. PinPoint publishes UK Declarations of Conformity on its own site; the gynaecological one (REC-777, 13 May 2026) names PinPoint_Gynae v1.2. The article and open question 3 have been rewritten; the finding that the MHRA register carries no version field is unchanged.
    Corrected 27 Jul 2026 — We wrote that the 578-woman ultrasound comparison appears on one page and is attributed to no source. It also appears in PinPoint's own press release of 8 July 2026, which attributes it to “PinPoint sub-analysis, data on file.” Our finding that no published study, dataset or statistical report stands behind it is unchanged and now confirmed by the company's own footnote.
    Corrected 27 Jul 2026 — We wrote that the 90,000 referral denominator cites nothing. The same press release footnotes it — to hysteroscopy volumes scaled for population growth and to NHS Scotland guidance, not to any count of post-menopausal-bleeding referrals. The chain step now says so and is tagged a proxy figure rather than unsourced. Corrections log →