AI for Science

Oxford AI Predicts Heart Failure Five Years From CT

Oxford medicine researchers trained an AI on 70,000+ CT scans to flag heart failure risk five years early at 86% accuracy; top-risk patients are 20x more likely to develop it.

Oxford AI Predicts Heart Failure Five Years From CT — article cover

In early April 2026, the University of Oxford’s Radcliffe Department of Medicine announced a new tool: an AI that reads routine CT scans and predicts whether a patient will develop heart failure at least five years before it happens. The model was trained on anonymized CT scans from more than 70,000 patients, predicts with roughly 86% accuracy, and flags a highest-risk group that is around 20 times more likely to develop the condition than everyone else.

The value of this kind of work is not “one more AI classifier.” It is the choice of target: a condition that is notoriously hard to catch early and devastating when caught late. By the time heart failure is diagnosed, substantial damage is usually done. A risk signal five years ahead is exactly the raw material preventive medicine needs.

The Model Reads Fat, Not the Heart Itself

The most interesting detail is where the signal comes from. The model does not measure the geometry of heart chambers or vessels; it detects textural changes in the fat surrounding the heart. Fat has long been treated as bystander tissue in medical imaging, skipped by most radiology reads. This research turns it into a carrier of prognostic information. The practical implication: the model’s inputs already exist inside every chest CT — no new scanning protocol, no additional contrast agent. In principle, any chest CT could emit a heart-failure risk score as a side effect.

This tracks a broader pattern in AI medical imaging: the highest-value findings are less about seeing more accurately than a radiologist and more about seeing where nobody looks.

Why “Five Years” Is the Number That Matters

Heart failure is a progressive condition. Between the heart starting to decompensate and obvious symptoms appearing, there is a window where medication and lifestyle intervention can still change the trajectory for many patients. The problem is that this window has almost no reliable screening today — by the time a patient walks in with breathlessness or edema, half the window is gone. A tool that marks high-risk groups five years out at 86% accuracy pulls the intervention point forward. Combined with the 20x risk stratification, it is not replacing diagnosis; it answers the screening question of who deserves priority follow-up, which is precisely the question health systems are worst at answering today.

The Distance Between the Study and the Ward

Real hurdles remain before this reaches routine care. First, the training data comes from existing CT archives, which means the model inherits whatever population biases were baked in when those scans were collected; performance outside the training population needs prospective validation. Second, “86% accurate” is a population-level number. At the level of individual decisions, the costs of errors are asymmetric: labeling a low-risk patient as high-risk costs anxiety and follow-up, while missing a truly high-risk patient costs the prevention window entirely. Where to set the threshold is ultimately a clinical and policy question, not just a modeling one. Third, routine CT is not a risk-free universal screen — radiation dose and cost both argue against mass scanning. The realistic deployment path is the free-rider one: any patient who gets a chest CT for another reason also receives a heart-failure risk score.

What It Suggests for AI Health Products

Two takeaways. First, silent biomarkers are the largest untapped seam in AI medical imaging — signals already present in existing exams, with vast historical datasets that make retrospective validation far cheaper than collecting new data. A model that finds prognostic value in tissue every radiologist currently skips gets its training corpus essentially for free, and can be audited against years of outcomes that already exist. Second, risk-stratification products should position themselves as schedulers, not diagnosticians: they tell the health system who needs earlier, denser follow-up. Under that framing, both the regulatory path and the liability boundary are considerably cleaner than for products that issue diagnoses outright. The Oxford tool is a textbook case of both principles — an ordinary scan, an ignored tissue, and an output that ranks patients rather than names a disease.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL