AI for Science

NHS Trial: AI Breast Screening Detects 10.4% More Cancers

A University of Glasgow team published GEMINI in Nature Cancer: Mia AI in NHS Grampian breast screening lifted detection 10.4%, cut reading workload over 30%, and cut notification to 3 days.

NHS Trial: AI Breast Screening Detects 10.4% More Cancers — article cover
On this page6 SECTIONS
  1. How the GEMINI Evaluation Worked
  2. Three Numbers That Matter
  3. Who Gets Replaced — and Who Doesn’t
  4. From Pilot to the EDITH Trial
  5. Lessons for Health AI Teams
  6. Sources

On March 13, 2026, a University of Glasgow team published the GEMINI study in Nature Cancer showing that integrating the AI tool Mia into NHS Grampian’s breast screening program increased cancer detection by 10.4% while potentially cutting radiology reading workload by more than 30%. This is the UK’s first comprehensive evaluation of AI in breast cancer screening — and it deliberately did not stop at a single number. The team tested seventeen different AI integration scenarios, simulating where AI might sit inside the real clinical workflow.

That design choice separates GEMINI from the usual “AI versus doctor” studies. The question it asks is the one hospitals actually face at procurement: not whether to use AI, but where to put it and who signs off. For any team deploying clinical AI, that is the more honest form of evidence.

How the GEMINI Evaluation Worked

The evaluated tool is Mia, developed by Kheiron Medical Technologies (now part of DeepHealth Inc.), tested inside NHS Grampian’s screening program. The study covers screening data from 10,889 women, with computer simulations reconstructing how AI would behave once deployed — an approach the researchers describe as testing multiple workflow configurations that had never been tried before in this field.

The baseline numbers explain why this matters. Under the current UK system, two radiologists independently read each mammogram, yet roughly 20% of cancers are still missed. At the same time, only about one in five recalled women turns out to have cancer, so the recall mechanism is under pressure from both misses and false alarms. Mia was slotted into that structure rather than replacing it.

Three Numbers That Matter

  • Detection up 10.4%: the additional cancers found on top of the existing two-reader process
  • Workload down more than 30%: the reading effort saved in the best scenarios
  • Notification cut from 14 days to 3: how long affected women wait for results

The third number gets the least attention and matters most to patients. Compressing a two-week wait for life-changing news into three days is a capability upgrade for the screening program itself, not just a cost saving.

Who Gets Replaced — and Who Doesn’t

The study’s most important finding is about placement. The two best-performing configurations used AI either as the second reader (replacing one human reader) or as an additional safeguard reader — and both improved early detection without increasing recall rates. In other words, the question was never whether AI can replace radiologists. It is that once AI takes over the repetitive first or second read, scarce specialist time can be reallocated to harder interpretations and patient communication.

This is also a key piece of evidence shifting medical AI narratives from “replacement” to “redistribution.” Raising detection without raising recalls means the model’s errors skew toward catching more, not false-alarming more — exactly the failure mode a screening program will pay for.

From Pilot to the EDITH Trial

The funding and institutional context matters as much as the results. GEMINI was funded by the NHS AI in Health and Care Award in partnership with the National Institute for Health and Care Research (NIHR), with lead author Professor Gerald Lip supported by the Scottish Government’s Chief Scientist Office Innovation Fellowship. The University of Glasgow states plainly that this evidence fills the gap flagged by the UK National Screening Committee and supports the upcoming UK-wide EDITH trial.

Moving from one region’s seventeen-scenario simulation to a national multi-center trial is the standard British path from medical-AI paper to routine service: prove the placement first, then scale.

Lessons for Health AI Teams

Three actionable conclusions. First, experimental design around deployment position beats model scores: the seventeen-scenario simulation framework is worth borrowing directly by any clinical AI team. Second, report the metrics clinical managers actually run on — detection rate, recall rate, workload, notification time — not just AUC. Third, the funding channels have matured: programs like the NHS AI in Health and Care Award paying for workflow evaluation means clinical validation is now treated as part of product development, not as a post-publication afterthought. One more detail deserves attention: fairness is named in the paper’s own title alongside accuracy and clinical implementation, which means the team examined how the tool performs across different population groups — a step no national-scale screening deployment can skip. The next phase of AI in medicine will be won by the patience to embed models into institutions, one workflow at a time.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL