Six protein-based estimates of biological age moving in the same direction would normally sound like a straightforward result. A September 7 paper in Nature Biotechnology makes that result interesting precisely because its meaning remains unsettled. Researchers studying rentosertib, an experimental drug for idiopathic pulmonary fibrosis, found lower predicted biological ages across six clocks. The finding concerns measurements in blood; it does not establish longer life or preserved healthspan.
TENS Magazine’s analysis is that the study should be judged on two separate questions: whether the instruments detect a treatment response, and whether that response can stand in for a benefit people experience. Agreement strengthens the first case. It cannot, by itself, settle the second.
Agreement is a measurement result
The analysis used serum samples from 42 participants in a 12-week randomized, double-blind, placebo-controlled phase 2a trial. The underlying trial enrolled 71 people; the biomarker group included those with consent and complete sampling. These are human observations interpreted through computational models, not a trial demonstrating rejuvenation in healthy adults.
The paper reports 21 statistically significant treatment-versus-placebo comparisons among 54 clock, timepoint and regimen combinations, using a false-discovery threshold below 0.10. The strongest convergence occurred at week four. Thus, agreement across the six models should not be read as every clock passing every comparison throughout treatment.
For an editorial assessment, the useful distinction is between breadth and replication. Multiple models provide breadth: the result is less dependent on choosing one favored instrument. Independent replication would require new participants and another test of the hypothesis. Running additional models on the same specimens does not create additional clinical trials.
Prediction and intervention answer different questions
Why take protein clocks seriously at all? A 2024 Nature Medicine study led by M. Austin Argentieri developed a clock using UK Biobank data from 45,441 people. It identified 204 proteins for age prediction and tested the model in additional populations in China and Finland. Protein-predicted age was associated with future chronic disease and mortality, as well as measures of physical and cognitive function.
That work gives the measurement an empirical foundation beyond a neat numerical fit to birthdays. But it was a population study, not an experiment establishing that changing the score changes someone’s future. Its geographical validation supports a particular kind of portability; it does not automatically validate every use of the clock in a drug trial.
Putting the two studies together reveals an incomplete bridge. One connects a protein profile with subsequent risk. The other shows profiles responding during treatment. The missing connection is whether a treatment-induced shift reliably forecasts better outcomes. TENS’s assessment is that the field needs evidence for that connecting step, rather than treating the two existing findings as interchangeable.
Disease response complicates the interpretation
The rentosertib authors explicitly acknowledge that their clocks cannot separate effects on fibrosis from effects on aging. Their pathway analyses explore that distinction indirectly. The modest sample, short observation period and lack of complementary molecular measurements limit interpretation. Several authors work for Insilico Medicine, the drug’s developer, including its chief executive.
This creates a demanding test of interpretation. If a disease-related protein signal improves, a model may return a younger estimate even without evidence of a broader change in aging. That possibility does not make the measurement worthless. It changes the claim the measurement can support, and it makes independent confirmation particularly valuable.
A useful comparison would therefore keep disease-specific outcomes beside biomarker changes, rather than allowing the age estimate to replace them. If the two diverge, the divergence is information to investigate. A favorable clock reading should not erase an unfavorable functional result, and a useful disease response should not need a rejuvenation label to matter.
The standard is benefit, not a younger number
The FDA’s explanation of biomarkers and surrogate endpoints provides a separate framework. A biomarker can indicate a biological response. A clinical outcome measures an effect on how people feel, function or survive. Substituting one for the other requires supporting evidence; even validated substitutes can miss effects relevant to a treatment’s overall benefits and risks.
Applied here, that framework suggests a practical editorial scorecard for follow-up research: specify the population, declare the analysis in advance, reproduce the response, and test its relationship to meaningful outcomes. These are proposed standards for assessing the evidence, not claims that the current experiment has already met them. Longer observation would also help distinguish a persistent signal from a short-lived change.
The valuable prospect is a trial that learns more from each participant while keeping its claims proportionate to what was measured. Rentosertib’s clock results provide a reason to investigate that prospect. They offer no basis for recommending the drug for longevity, and no demonstrated number of healthy years gained. The next decisive result would connect a reproducible biomarker change to a benefit beyond the model’s output.
TENS Magazine conceptual illustration

