A biological-age score can look precise while measuring people unevenly. New human research in eLife reports that several widely used DNA methylation clocks lost accuracy in genetically admixed cohorts, especially among participants with substantial African ancestry. The finding does not show that any group ages faster or slower. It shows that a model trained in one population may misread another.
That distinction matters as epigenetic clocks move from research papers into conversations about health risk, environmental stress and longevity. The study, led by researchers at the University of California, San Francisco, the University of Miami and Case Western Reserve University, turns population diversity from a demographic footnote into a measurement-quality test.
What the researchers tested
The team applied first-, second- and third-generation methylation clocks to blood samples from 621 people in the MAGENTA Alzheimer’s disease study. Participants came from White, African American, Puerto Rican, Cuban and Peruvian cohorts and included Alzheimer’s cases and matched controls. The researchers then examined more than 2,500 people in three independent replication datasets: two African American cohorts and one White Swedish cohort.
These clocks estimate age or aging-related traits from methylation, chemical tags attached at selected CpG sites in DNA. In MAGENTA controls, the widely used Horvath clock’s correlation with chronological age was 0.72 in the White cohort, compared with 0.51 among African Americans and 0.45 among Puerto Ricans. The Cuban and Peruvian estimates were closer to the White result, although those MAGENTA subgroups were small.
The pattern persisted in replication. Across the full age range, the Horvath clock’s correlation with chronological age was 0.97 in the Swedish cohort, 0.85 in the GENOA African American cohort and 0.88 in the Grady Trauma Project cohort. Error was larger in the latter groups. Other clocks varied, but cohorts with more African ancestry were repeatedly among those with the lowest accuracy.
A model error is not a biological verdict
The most important editorial conclusion is also the easiest to miss: a difference in a clock output cannot be interpreted as a difference in aging until the clock itself performs comparably across the populations being compared. Without that check, apparent “age acceleration” may partly reflect transport failure in the algorithm.
The Alzheimer’s analysis makes that risk concrete. Most clocks did not consistently separate cases from controls in the admixed cohorts. DunedinPACE showed a more consistent signal in White, African American and Puerto Rican groups, but not in Cuban or Peruvian groups. Even that result is an association in existing samples, not evidence that the measure diagnoses Alzheimer’s disease or predicts an individual’s future.
To explore why performance shifted, the authors examined the CpG sites used by the clocks. Nearly one-quarter of sites in both the Horvath and PhenoAge clocks were differentially methylated between African- and European-ancestry reference groups. The team also found that genetic variants influencing methylation at clock sites, known as methylation quantitative trait loci, were more frequent in African-ancestry data. Those variants may contribute to error, but they did not explain it completely.
The portability ladder longevity tools need
A clock should therefore clear a portability ladder before its number is treated as actionable: comparable calibration across populations, replication in independent cohorts, stable performance within relevant age ranges, and added value beyond established clinical and social predictors. Passing only the first step—correlation with age in one dataset—is not enough.
This is where the new paper and earlier National Institute on Aging guidance converge. NIA has described epigenetic clocks as promising biomarkers, while also noting that demographic, socioeconomic, behavioral and mental-health factors can equal or outperform them for some late-life outcomes. The institute’s biomarker program emphasizes validation across human populations and longitudinal cohorts. The new study shows why that validation must include model portability, not simply a larger sample.
The practical future is not a different clock for every identity label. It is transparent training data, ancestry-aware evaluation, environmental information, and models designed around CpG sites that are less vulnerable to population-specific genetic effects. A useful tool should also report uncertainty instead of compressing biological complexity into a falsely universal age.
What remains uncertain
This was a computational and observational analysis of existing human blood samples, not a treatment trial. It measured model performance and associations; it did not demonstrate a cause of aging, change healthspan or lifespan, or validate a consumer test. MAGENTA was relatively small, did not represent every ancestry combination and used blood to study a brain disorder. Environmental exposures were not comprehensively available, and methods for estimating blood-cell composition may themselves generalize unevenly.
The study also did not test every methylation clock. Its genetic analyses could identify contributors to error without proving that ancestry-linked variants caused all performance differences. Larger longitudinal cohorts, clearer reporting of training populations and prospective validation will be needed before a methylation score can be trusted equally across individuals.
The paper’s lasting contribution is a standard, not a score. Longevity biomarkers should be judged not only by how impressively they predict an average, but by whose biology they measure accurately—and how clearly they disclose when they do not.
TENS Magazine conceptual illustration


