Skip to content
Longevity

Sex-Specific Aging Clocks Put the Reference Population First

A new study makes reference populations central to aging-clock interpretation. TENS proposes clearer reporting of model versions, biomarker changes and clinical claims.

Share Email
Conceptual DNA strand and timing rings representing aging measurement; not a study image
TENS Magazine conceptual illustration

A biological-age score needs a reference population before it can tell a useful story. A Nature Medicine paper published September 16 brings that choice into focus by building sex-specific aging clocks. For longevity research, the immediate question is how a score should be interpreted when the people used to define its baseline change.

The MULTI Consortium developed 38 clocks across 15 organ systems. Its human-data analyses found sex-dependent molecular patterns and associations with disease and mortality. The authors also emphasize that pooled models remain useful. Separate models are a way to investigate differences, not a universal replacement for shared models.

TENS Magazine’s analysis is that aging-score reports should separate three decisions: the reference used to calculate a deviation, the outcome that deviation predicts, and the evidence required to interpret a change. Calling all three “biological age” makes a compact label carry more information than readers can reasonably recover from a number alone.

The reference belongs beside the result

The new study’s limitations matter here. Genetic analyses were restricted to European ancestry; sex was modeled as binary and genetically proxied. Hormonal status and gendered exposures were not explicitly incorporated. Repeat measurements were limited, and external validation at comparable scale remains necessary.

Those boundaries suggest an editorial reporting rule: put the comparison population beside the headline result. A score should arrive with an explanation of whose data established its reference. Readers should not have to search a methods supplement to discover whether a seemingly universal measurement was developed within a much narrower group.

Consider a hypothetical report that changes after its reference model is updated, while the underlying sample stays identical. That difference would describe a changed interpretation of the sample. It would not, by itself, document a biological change in the person. This is an illustrative distinction, not an outcome reported by the new study.

A useful disclosure would therefore name the model version as well as the reference group. It would preserve the distinction between a new reading of old information and new information about someone’s health. That is a practical requirement for understandable research communication, even before any clinical application is considered.

Prediction and responsiveness ask different questions

An earlier Nature Medicine analysis supplies a complementary test. Published August 21, the TranslAGE study assembled 51 longitudinal human intervention studies and calculated 16 epigenetic clocks alongside 94 other DNA-methylation biomarkers. It examined how measurements responded across interventions rather than treating every clock as interchangeable.

The investigators reported that population characteristics and study duration influenced responsiveness. They also identified unresolved questions about linking short-term clock changes to longer-term health outcomes. Differences in study quality and incomplete harmonization of preprocessing limit comparisons. These are human biomarker data, not proof that a particular clock reduction delivers longer life.

Read alongside today’s paper, that work supports a reporting distinction between “different from a reference” and “different after an intervention.” The first comparison can exist without repeated sampling. The second needs observations over time. Neither phrase, standing alone, tells a reader what happened to disability, disease or survival.

TENS would keep those statements in separate sentences in a research report. A baseline score, a subsequent movement and a clinical outcome should each retain its own denominator, time frame and uncertainty. Combining them into one rejuvenation claim would erase the very distinctions that make the underlying studies informative.

A benefit claim needs its own evidence

The US Food and Drug Administration draws a further distinction between a biomarker and a surrogate endpoint. A biomarker measures a biological characteristic; a validated surrogate has evidence supporting its use in place of a clinical outcome in a particular setting. FDA also cautions that surrogate measures can miss parts of a treatment’s overall benefits and risks.

For an aging-clock report, the corresponding editorial question is precise: what evidence connects this change, in this setting, with an outcome that matters to people? An answer about prediction does not automatically answer a question about treatment response. An answer about response does not automatically establish the balance of benefit and harm.

The resulting TENS proposal is a three-part score description: reference, observed behavior and permitted interpretation. Each part should be stated independently, and an unanswered part should remain visibly unanswered. This format would let researchers explain an advance without making an early measurement carry the authority of a completed clinical trial.

Today’s study makes the reference question harder to ignore. The opportunity for longevity infrastructure is to make that complexity readable: preserve the population definition, distinguish a changed model from a changed person, and keep clinical claims tied to clinical evidence. A more informative score can advance research while leaving lifespan and healthspan benefits unproven.

TENS Magazine conceptual illustration