Skip to content
Longevity

A Retinal-Age AI Finds Risk Signals, Not a Longevity Score

A large human imaging study links an AI-derived retinal age gap to later health risks, but calibration drift and observational evidence keep it from being a clinical longevity score.

Share Email
Conceptual illustration of a retina-inspired vascular network and computational aging signals
TENS Magazine conceptual illustration

A color photograph of the back of the eye may carry more information about aging than its clinical familiarity suggests. A peer-reviewed Nature Communications study published August 26 reports that a retinal foundation model estimated chronological age with an average error of 2.85 years in 71,343 UK Biobank participants. The difference between the estimate and a participant’s actual age—the retinal age gap—also tracked a wide set of health measures and later outcomes.

The result is substantial because the retina offers a noninvasive view of neural tissue and small blood vessels. It is also easy to overread. The model was trained to predict calendar age, not to measure a universal biological-age substance. Its risk findings are observational associations, not evidence that changing a retinal score would prevent disease, extend healthspan or lengthen life.

What the model actually learned

The researchers fine-tuned RETFound, a foundation model originally developed from large collections of unlabeled retinal images. After quality control, they analyzed 130,360 fundus images. Predictions were generated out of sample through five-fold cross-validation, with images from the same person kept together to avoid leakage between training and testing sets.

Averaging available images from each participant improved performance from a 2.99-year image-level error to 2.85 years at the participant level. The team then tested age prediction in 4,757 people from the independent Rotterdam Study. Error increased to 3.35 years, and the model systematically underestimated age there, particularly in men. Differences in cameras, image acquisition, health or age distribution could all contribute to that shift.

The study separates two validation problems that are often bundled together. The Rotterdam analysis shows that the model can still estimate age in a second cohort, although with measurable calibration drift. The reported links to future disease and death, however, came from the UK Biobank analysis. Transporting the clock is therefore not the same as transporting every clinical association attached to it.

A risk signal is not an aging intervention

Within UK Biobank, a higher retinal age gap was associated with measures spanning metabolism, inflammation, smoking, cardiovascular fitness, cognition and eye health. In roughly 15 years of follow-up, the gap was also associated with all-cause mortality and incident cardiovascular disease, stroke, dementia, cancer and several other outcomes after adjustment for demographic and conventional risk factors. Reported hazard ratios ranged from 1.03 to 1.20 per standard-deviation increase in the retinal age gap.

Those prospective associations strengthen the case that the images contain information related to health. They do not reveal a single retinal mechanism driving every outcome. Blood vessels, the optic disc, the macular region, pigmentation, image intensity and other features can all influence a model’s representation. Some may reflect aging biology; others may reflect disease, anatomy, equipment or population differences.

That matters because a clock can be accurate at reconstructing age yet remain uncertain as a decision tool. The practical question is not whether a score correlates with many outcomes in a large dataset. It is whether the score adds stable, actionable information beyond established assessment in a new clinic, on a new camera and across populations that differ from the training cohort.

Sex differences sharpen the calibration question

Female- and male-specific models achieved similar age-prediction accuracy, but the patterns associated with their retinal age gaps differed. In men, links to metabolic syndrome were stronger. In women, the model’s attention and genetic findings pointed more toward retinal vasculature. Retinal aging patterns also varied across the menopausal transition.

These observations should not be converted into separate health instructions. The investigators noted that body and eye size could influence apparent sex differences, and adjustment for height reduced but did not eliminate the gap. Sex-stratified dementia and cancer analyses also had limited event counts and wide confidence intervals. The study is better read as a warning against assuming that one pooled score has identical meaning in every group.

The next evidence standard

Both UK Biobank and Rotterdam were predominantly composed of people of White European ancestry and represented relatively healthy aging populations. The researchers also excluded the lowest-quality quarter of images. That improves internal analysis but leaves an important deployment question: how would the clock perform in more diverse clinical populations, in people with substantial eye disease and with the imperfect images produced in routine practice?

The next decisive test is a prospective, multi-site calibration study that freezes the model before enrollment, uses several camera systems, prespecifies performance across ancestry, sex and eye-disease groups, and compares the retinal score with standard risk tools. It should ask whether the score changes a defined clinical decision and improves an outcome, rather than merely finding another association after the fact.

The new work therefore advances retinal aging from a clever age-estimation exercise toward a richer human risk phenotype. Its strongest contribution is not a consumer-ready longevity number. It is a demonstration that retinal images contain layered signals whose technical portability, biological meaning and clinical usefulness must be validated separately.

TENS Magazine conceptual illustration