Skip to content
Longevity

Aging Risk Scores Need Validation That Matches the Claim

The new CKMAI study shows why aging-risk validation must specify populations, outcomes and remaining limits.

Share Email
Conceptual DNA strand and timing rings representing aging measurement; not a study image
TENS Magazine conceptual illustration

A new aging score should arrive with a clear account of the future it is trying to predict. Death, a first cardiovascular event and an existing disease classification are different outcomes. Putting them under the same longevity label can make a promising research tool appear further along than its evidence supports.

Research published October 6 in PLOS Medicine introduces the cardiovascular-kidney-metabolic aging index, or CKMAI. Zhengyang Zhu and colleagues developed the index using 6,896 adults in the US National Health and Nutrition Examination Survey. Its 18 inputs combine clinical and demographic information. The researchers reported stronger prediction than several existing aging indices.

TENS Magazine’s analysis is that the important comparison is between validation tasks, not headline scores. Reading CKMAI alongside the American Heart Association’s PREVENT research suggests a practical standard: every claim of predictive progress should identify the population, outcome, time horizon and genuinely new test data behind it.

Match the question before comparing the score

The original PREVENT study, led by Sadiya Khan and published in Circulation, addressed incident cardiovascular disease in adults aged 30–79 without known cardiovascular disease. Development and validation together included more than 6.6 million people. External validation used 21 additional datasets and assessed both discrimination and calibration.

Discrimination concerns how well a model separates people with different outcomes. Calibration concerns whether estimated risks agree with observed risks. A system can rank people usefully while assigning probabilities that need correction. Both questions matter when a number is expected to inform a real decision.

CKMAI’s authors explicitly state that they did not perform a formal quantitative comparison with PREVENT. Their outcomes include mortality and high-risk CKM status, rather than the same incident-event task.

Our interpretation is that declaring a winner from the two papers’ performance numbers would answer a question neither comparison establishes. A fair head-to-head evaluation would first align eligibility, endpoints and follow-up. Otherwise, differences in the difficulty of the prediction problem can be mistaken for differences in the quality of the model.

This distinction also protects useful specialization. An index may help characterize one population without needing to replace every established risk equation. The editorial question is what additional uncertainty it resolves for a specified use, and whether that contribution survives a comparison designed for that use.

External validation needs an outcome attached

The CKMAI study included a separate Chinese hospital cohort of 261 patients. That cohort lacked mortality follow-up; it tested identification of high-risk CKM status. The study also tested performance across different NHANES time periods.

Those checks provide different kinds of evidence. Testing across time can challenge a model with changes in the population or measurement setting. Testing at another institution can challenge its transportability. But a successful test for existing disease status cannot, on its own, establish accurate prediction of later deaths.

TENS would therefore describe validation as a map with labeled destinations. Each dataset should be paired with the outcome it actually observed. This prevents a broad phrase such as “externally validated” from lending the authority of one test to a different, unfinished test.

The PREVENT work offers a useful reporting comparison because it identifies additional datasets and evaluates observed cardiovascular events. That does not make its results transferable to CKMAI. It shows why the endpoint belongs next to the validation claim, where readers can inspect both together.

An aging label does not establish reversibility

CKMAI’s observational design cannot establish causation. Its authors identify possible selection bias, residual confounding and single-time-point biomarkers among the limitations. Larger, diverse external cohorts with mortality outcomes remain necessary.

For longevity research, there is another question beyond prediction: what would count as evidence that changing the score changes a person’s future? A lower number after an intervention could reflect a changed input without proving slower aging. Establishing benefit requires measuring meaningful outcomes under a design capable of testing that claim.

This is where a research roadmap should keep separate milestones. First comes reproducible calculation. Next comes reliable prediction in the intended population. After that comes evidence that using the prediction improves decisions and outcomes. Completing an earlier milestone is progress, but it cannot silently stand in for the later ones.

Our proposed reporting standard would make an aging index’s evidence easy to audit: state who was studied, what was predicted, where that prediction was independently checked and what remains untested. For a health system considering future evaluation, those details would be more informative than a single impressive accuracy figure.

CKMAI is human observational research, not a trial demonstrating longer life or improved healthspan. Its contribution deserves examination at the level its design supports. The next advance would be evidence that travels across settings and retains a precise connection to the outcome being claimed.

TENS Magazine conceptual illustration