Epigenetic Clocks
Epigenetic Clocks
Why the Number Everyone Uses Is the Least Reliable One
A methylation age arrives with units. Someone is told their epigenetic age is 47.3 while their passport says 42, and the four-year gap reads like a measurement — a quantity that exists in the body and was determined by an instrument. It is a regression prediction, and the gap is a residual from that regression.
Two earlier pieces in this series dealt with the layer underneath: whether a bisulfite conversion or a basecalling model gets the methylation state of a given cytosine right. This one grants that layer entirely. Assume every beta value is correct. The question here is what happens between a correct methylation profile and the number that reaches a paper, and the answer is that three transformations sit in between, each of which adds something the reader cannot see.
Start with the probes
Individual CpG measurements are noisier than most people using them assume. In the study that quantified this most directly, low-reliability clock CpGs tended to have extreme values near 0 or 1 and low variance, and while CpGs with strong associations with age or mortality tended to have higher reliability, the authors note those associations are themselves artificially depressed by technical noise1.
That is a circular trap worth naming: a noisy probe looks less associated with age because it is noisy, so probe selection quietly favours reliable probes, and the reliability of the ones that got in still is not uniform.
What that does to the clock
Run the same DNA twice and the clocks return different ages. Across six prominent clocks, technical noise produced deviations of up to nine years between replicates2. Broken down: the widely used Horvath multi-tissue clock showed a median difference of 2.1 years between technical replicates and a maximum deviation of 5.4 years; across the other clocks, median deviation ranged from 0.9 to 2.4 years and maximum difference from 4.5 to 8.6 years1.
Two details make this worse than it first sounds.
First, the obvious remedy fails. Eliminating low-reliability CpGs does not ameliorate the issue2 — the noise is distributed across the probe set rather than concentrated in a removable subset.
Second, the discrepancies were largely uncorrelated with each other and with age and sex, and there was little to no systematic batch effect for any clock except one1. That rules out the comfortable explanation. This is not a batch effect you could model out. It is noise, and noise does not cancel when you compare two people.
The step that matters, and the number that breaks
Here is the part I would put in front of anyone about to design a study around this.
The reliability of the clocks themselves looks excellent by conventional standards: intraclass correlations ranging from 0.917 to 0.979, with the Horvath multi-tissue clock at 0.9451. An ICC above 0.9 is what you would want from any assay.
But almost nobody analyses predicted age. What gets used is age acceleration — the deviation after adjusting for chronological age — because that is the quantity that could reflect biology rather than the calendar. And the authors state the consequence plainly: age acceleration has lower ICCs because of reduced biological variance, ranging from 0.755 to 0.948, with the Horvath clock falling to 0.8171.
An ICC is a ratio of true variance to total variance. Predicting age across an adult cohort spans decades, so the denominator is enormous and the noise looks small against it. Subtract chronological age and the spread collapses to a few years, while the technical noise stays exactly where it was. The same measurement error, divided by a much smaller variance, produces a much worse reliability coefficient.
Which means the reassuring statistic and the reported quantity are not the same object. A paper can honestly describe a clock with an ICC of 0.945 and then analyse a derived variable whose ICC is 0.817 — a number that, on the same sample run twice, moves by a year or two for reasons that have nothing to do with the person.
A residual is defined relative to a crowd
There is a second property of age acceleration that is easy to miss because it is a matter of definition rather than of noise.
Age acceleration is calculated by taking the residual of the clock values regressed on age3. The regression line is fitted within the sample being analysed. So a person's age acceleration is not a property of that person — it is a statement about where they sit relative to the age-methylation trend in this particular cohort.
Recruit a healthier comparison group and the line moves; the same individual's value changes without anything about them changing. This is not hypothetical: work on a long-lived population found that calibrating against a regional reference group altered the estimated aging advantage, reducing an initial three-year estimate to roughly two years and improving consistency across clocks4.
The practical consequence is that age acceleration values are not portable between studies in the way a haemoglobin concentration is. Two papers reporting “two years of age acceleration” may be reporting residuals from two different lines.
The clocks are measuring different things, by construction
It is common to see several clocks run on one dataset and treated as replicate attempts at a single quantity. They are not.
First-generation clocks such as Horvath and Hannum were trained to predict chronological age. Second-generation clocks such as PhenoAge and GrimAge were trained on clinical markers and mortality. Third-generation measures such as DunedinPACE estimate a rate of aging rather than an age at all3. Different training targets, different CpG sets, different outputs.
So disagreement between them is not a malfunction — it is what should be expected from estimators of different quantities. A systematic review found the major clocks tend to agree in the direction of effects but vary in size, and that chronological and mortality-trained clocks diverge more for psychiatric than for physical outcomes5. Agreement across clocks is genuine corroboration precisely because they are not the same instrument; the error is in treating a disagreement as one of them being wrong.
Where the associations hold, and where they stop
Being fair to the field matters here, because the clocks do carry real signal. In a nationally representative US cohort followed for nearly two decades, overall mortality was predicted most strongly by GrimAge acceleration, followed by Hannum, PhenoAge, Horvath and Vidal-Bralo6. These are not empty measures.
The same study also reports the limit, and it is the finding I would keep. Overall mortality prediction differed by race and ethnicity for several clocks — and despite being predictive in non-Hispanic White participants, Horvath, Hannum and GrimAge age acceleration failed to predict overall mortality in Hispanic participants6.
That is the same structure as the portability problems covered elsewhere in this series: a model fitted in one population, applied in another, producing a number that is computable everywhere and predictive only where it was built. The number does not announce which case it is in.
| Quantity | What it actually is | Reliability on a repeat run |
|---|---|---|
| Beta value at one CpG | A measurement, subject to assay noise | Varies widely by probe1 |
| Predicted epigenetic age | A weighted sum fitted to a training cohort | ICC 0.917–0.9791 |
| Age acceleration | A residual, relative to this cohort's trend line | ICC 0.755–0.9481 |
| Change in age acceleration | A difference between two residuals | Noise compounds across both2 |
| PC-based versions | The same clocks refitted on principal components | Replicates within ~1.5 years2 |
The fourth row is the one that matters for intervention studies, and it is the worst case: an intervention effect is a difference between two quantities that each already carry a year or more of technical variation. The fifth row is the field's own fix, discussed below.
The field's response
This is not a problem being ignored. The same group that quantified the noise proposed a solution: compute principal components from the CpG-level data and predict biological age from the PCs rather than from individual probes, which extracts shared age-related variation while diluting probe-specific noise. The retrained PC versions of six clocks bring most technical replicates within about 1.5 years, with equivalent or improved prediction of outcomes and more stable longitudinal trajectories2.
That is a substantial improvement, achieved with one extra step and no need for replicates during training2. Independent evaluation supports it — PC-based and multi-system measures show the highest technical reproducibility across array types, slide positions and DNA extraction protocols7.
Two caveats keep this from closing the question. Improving technical reliability does not make a residual portable across cohorts, and it does not turn a prediction into a measurement — the PC clocks are more stable estimators of the same constructed quantity. And the same independent evaluation reports that short-term exposures reduce the biological reliability of clocks, with the authors suggesting that biological rather than technical variation accounts for much of the instability seen in downstream applications7. Fixing the noise reveals how much movement was never noise.
What to do about it
- Use PC-based clocks where reliability matters — longitudinal designs and intervention trials especially2. This is the cheapest available improvement and it requires one additional processing step.
- Include technical replicates and report the observed spread. A study measuring a two-year effect on a quantity with a two-year replicate deviation needs to show it has more precision than the literature default.
- Report the ICC of the quantity you analysed, not the clock. Those are different numbers, and the one you analysed is the lower of the two1.
- State how age acceleration was computed and in which sample the regression was fitted. Residuals from different lines are different quantities carrying the same name.
- Do not compare age acceleration values across studies numerically. Compare directions and effect sizes relative to each study's own reference, or recompute from raw data under one model.
- Choose the clock generation from the question. A chronological-age clock and a mortality-trained clock answer different questions; running both and reporting whichever agreed with the hypothesis is a search over outcomes.
- Treat cross-clock agreement as the corroboration it is, and treat disagreement as information about which construct is moving rather than as one clock being wrong5.
- Check whether the clock has been validated in the population you are applying it to. Predictive in one group and null in another is a documented outcome, not a hypothetical6.
- Never report an individual's epigenetic age as a finding about that person. On a repeat run of the same sample it would come back different, and the number was never calibrated to be read one person at a time.
The shape of the error
The recurring shape in this series is a computation that is correct and an interpretation attached to it that the computation does not support. Epigenetic clocks have a distinctive version, because the interpretation is embedded in the unit.
Years are a measurement unit. Attaching them to a regression output invites everyone downstream to treat the output as a measurement — something the body has, which an assay determined, and which should therefore be stable, comparable and personal. None of those three follow. It is not stable, because a repeat run moves it. It is not comparable, because the residual is defined against a cohort. And it is not personal in the way the unit implies, because the model was fitted on populations and validated on populations.
None of which makes the clocks useless. They predict mortality in large cohorts, they respond to interventions, and the reliability problem has an available fix. The failure is not in the tool but in the compression: a fitted value minus a fitted line, reported in years, read as an age.
The methylation is real. The years are a unit inherited from the training label, and everything the word “age” implies about permanence and personhood arrived with the unit rather than with the evidence.
References
- Higgins-Chen AT, Thrush KL, Wang Y, et al. A computational solution for bolstering reliability of epigenetic clocks. Detailed per-clock ICC and replicate-deviation figures cited here are from the openly available preprint version, bioRxiv 2021.04.16.440205 — verify against the published article before quoting. biorxiv.org — 2021.04.16.440205
- Higgins-Chen AT, Thrush KL, Wang Y, et al. A computational solution for bolstering reliability of epigenetic clocks: implications for clinical trials and longitudinal tracking. Nature Aging 2022;2(7):644–661. nature.com/articles/s43587-022-00248-2
- Epigenetic-based age acceleration in a representative sample of older Americans: associations with aging-related morbidity and mortality. PNAS 2023;120(9):e2215840120. Author list not captured — verify before publication. pnas.org/doi/10.1073/pnas.2215840120
- Calibrated comparison of epigenetic age estimates in a long-lived (blue zone) population against a regional reference group. Cited for the cohort-dependence of residual-based estimates; full citation to be confirmed before publication. researchgate.net — figure page carrying the summary
- A systematic review of biological, social and environmental factors associated with epigenetic clock acceleration. Ageing Research Reviews 2021. Author list not captured — verify before publication. sciencedirect.com — S1568163721000957
- Epigenetic age acceleration and mortality risk prediction in US adults (NHANES 1999–2002, n = 2,105, followed through 2019). Author list not captured — verify before publication. ncbi.nlm.nih.gov/pmc/articles/PMC12397028
- Biological versus technical reliability of epigenetic clocks and implications for disease prognosis and intervention response. bioRxiv 2025.10.13.682176 (preprint; not peer reviewed). biorxiv.org — 2025.10.13.682176

