Tests reporting a biological age produce a number that looks like a measurement of ageing. It is the output of a statistical model trained on a particular dataset to predict something else.
The most common approach uses chemical tags on DNA
Small chemical groups attach to DNA at specific sites and influence whether nearby genes are switched on, and the pattern changes with age in a reproducible way.
Researchers measure these tags at many sites, then fit a model selecting the subset that together best predicts a person's chronological age.
Applied to a new sample, the model returns a predicted age, and the difference between that prediction and actual age is reported as biological age.
The target defines what is captured
Early models were trained to predict chronological age, so by construction they captured what changes reliably with time rather than what causes decline.
A perfect model of that type would simply reproduce the date of birth and carry no additional information about health.
Later versions were trained instead to predict mortality or measures of physical function, and these correlate more strongly with health outcomes, because that is what they were built to track.
The output depends on the training population
A model learns relationships present in the data it was fitted to, including the age range, ancestry, health status and living conditions of that group.
Applied outside those conditions, the relationships may not hold, and different models applied to the same sample return different biological ages.
That disagreement is informative: it indicates the number reflects a modelling choice rather than an underlying quantity all the models are measuring.
Measurement noise is substantial
Results vary with the tissue sampled, since blood contains a mixture of cell types whose proportions shift with health and shift the reading with them.
Repeat measurement of the same person within a short period can produce differences of years, which limits how usefully small changes can be interpreted.
Reported reductions following an intervention need to be read against that variability before they are treated as real.
Where the field is useful
These measures are valuable in research, where averaging across large groups removes much of the noise and differences between populations become interpretable.
Their value to an individual buying a consumer test is far weaker, since a single noisy figure with no established action attached to it changes nothing.
Established clinical measurements remain the basis on which a doctor assesses individual health, and a biological age result does not substitute for them.