cfDNA Fragmentomics
cfDNA Fragmentomics
Why a Fragment-End Profile Records the Laboratory as Well as the Patient
Mutation-based liquid biopsy asks a question with a physical answer: is this base different from the reference. Fragmentomics asks something less concrete. It reads the shape of the molecules — how long they are, where along the genome the short ones concentrate, which nucleotides sit at their ends — and treats that shape as a cancer signal.
The shape is real biology. Plasma DNA is nucleosomal debris, cut by specific nucleases, and the cut positions record the chromatin state of the cells that died5. But shape is also the property that every step between the vein and the FASTQ file is capable of modifying. A tube additive, a delayed spin, a bead chemistry, an end-repair enzyme, a trimming parameter, and a line of code deciding where a fragment starts all act on the same quantity the biology is written in. The measurement contains both, and nothing in the output separates them.
This is a different problem from the one covered in the earlier piece on ctDNA detection limits. That piece was about how much tumour DNA is present and whether a given depth can see it — a sampling and limit-of-detection argument that holds even for a perfect assay. Here the amount is not the issue. The issue is that the signal being measured is a structural feature, and structural features are what laboratory procedures alter.
What is actually being measured
Three families of feature dominate the field. Fragment length: plasma cfDNA peaks near 167 bp with di- and tri-nucleosomal shoulders, and tumour-derived fragments run shorter, which is why size selection enriches the tumour signal3. Genome-wide fragmentation profiles: the ratio of short to long fragments computed in large genomic bins, which was the basis for detecting seven cancer types with sensitivities from 57% to 99% at 98% specificity1 and later for a prospective screening study of 958 individuals eligible for lung cancer screening, trained on 576 and validated on a held-out 3822. And end motifs: the first few bases at fragment termini, whose frequencies shift in cancer because the responsible nucleases have different sequence preferences45.
These are not equally fragile, and the ranking is the opposite of what intuition suggests.
Tube and time: less than folklore says, and not where you expect
The received wisdom is that cfDNA work lives or dies on the collection tube and the time to plasma, because leukocytes lyse and release long genomic DNA that dilutes the pool. That is true of yield. In 231 blood samples from 62 cancer patients, cfDNA levels rose steadily with time in K3EDTA tubes while remaining stable in stabilising tubes — yet sequencing showed negligible differences in background error or copy-number changes between the two9.
For fragmentomic shape specifically, the most direct experiment paired EDTA against two stabilising tubes in the same individuals and recorded processing times for a separate cohort. Neither the tube nor the delay significantly changed the fragment size distribution or the proportion of fragments between 100 and 150 bp6. Neither did they significantly change the diversity of fragment-end trinucleotides. A comparison of one-step against two-step centrifugation reached the same conclusion for nuclear cfDNA: no significant difference in size, in end-motif diversity score, or in genome distribution8.
Extraction sits between handling and library, and it does move the distribution. Comparing seven cfDNA extraction kits on the same material, both yield and recovered fragment size varied significantly between them10 — chemistries differ in how efficiently they retain short molecules, and short molecules are the enriched fraction. The same work found that across 219 clinical samples, fragments were shorter in plasma processed immediately after venipuncture than in archived samples, consistent with longer background DNA from lysed blood cells accumulating over time10. That is the leukocyte-lysis effect showing up where it belongs: as a change in what is in the tube, not as a change in the nuclease biology.
So the coarse features are sturdier than the folklore. What the same studies found instead is more uncomfortable.
First, the effects that did appear landed on the finer features. Genome-wide fragmentation patterns shifted with tube and, moderately, with processing delay6. Centrifugation protocol left nuclear cfDNA alone but changed the fragmentomics of mitochondrial cfDNA8 — a compartment with its own distinct fragmentation behaviour12, and one that contributes to genome-wide profiles.
Second, and more consequentially: when healthy control samples from two clinical centres were clustered on genome-wide fragmentation patterns and again on fragment-end trinucleotide frequencies, they separated by centre — despite having been processed under the same written protocol, and regardless of whether the blood was drawn into EDTA or a stabilising tube7. The authors drew the conclusion themselves: multi-centre studies using stabilising tubes can still carry a batch effect in fragmentomic analysis, and this extends to studies comparing cancer cases against healthy individuals collected at different clinical centres7.
That last sentence describes the design of a large share of fragmentomics discovery work.
One detail from the same study is worth carrying forward, because it shows how much the analysis decides. The proportion of short fragments differed significantly between male and female donors — but that difference did not appear in the genome-wide fragmentation profiles, because those profiles are z-score normalised across bins within each sample, which removes global shifts by construction7. The normalisation is doing exactly what it was designed to do. It is also deciding which effects, technical and biological alike, are still visible by the time anyone looks at the data.
The larger lever is the library, not the tube
The most systematic assessment to date took plasma from ten healthy donors, built libraries from each with nine different commercial kits, processed every library through ten trimming-and-alignment routes, and then repeated the feature extraction across 1,182 plasma samples from published studies11.
Kits differ in what they capture. One kit recovered a higher proportion of fragments in the 50–99 bp range; another recovered fewer short fragments and more in the 151–220 bp mononucleosomal band; three others recovered more di-nucleosomal fragments in the 300–380 bp range. One kit produced a median mitochondrial read fraction 4.4 times the median across all kits, consistently across every analysis route — an inherent property of its chemistry rather than a processing artefact11.
The clearest result is what happens in principal component space. After optimal processing, fragment length distributions from different kits converged reasonably well. End motifs did not: samples clustered by library kit11. The feature carrying much of the diagnostic weight in the current literature is also the feature that most reliably identifies which kit prepared the library.
The practical consequence is narrow and sharp. A model trained on end-motif frequencies from one chemistry and applied to samples prepared with another is being asked to generalise across a variable that explains a substantial share of the variance in its input — which is a different and harder task than the one its cross-validated performance estimated. The authors’ own recommendation follows from this: prefer features that are robust to experimental procedure, harmonise where the design permits, and choose models that tolerate batch structure11.
End repair writes some of the ends it reports
There is a specific mechanism behind part of that, and it deserves stating plainly because it is easy to miss in a methods section.
Standard double-stranded library preparation begins by blunting the input. The end-repair step removes 3′ protruding single-stranded ends and extends 3′ recessed ends, using the opposite strand as a template13. Plasma cfDNA arrives with a high frequency of exactly those irregular termini. So for a substantial share of molecules, the 3′ end recorded in the sequencing data is not the end the nuclease made — it is a sequence the polymerase synthesised in the tube. This is why end-motif work has concentrated on 5′ ends: the 3′ ends were not measurable under the standard protocol. Single-stranded library preparation, which ligates adapters directly to denatured molecules without end repair, preserves both native termini, and studies using it report diagnostic signal in the 3′ motifs that the conventional workflow had been destroying13.
The same choice reshapes the length distribution. Single-stranded preparation recovers short, nicked and single-stranded species that double-stranded protocols largely miss, revealing a 30–80 bp population and roughly ten times as many reads below 100 bases14. And because end repair replaces terminal bases with newly synthesised unmethylated nucleotides, it also depresses inferred CpG methylation near fragment ends, with measurable consequences for tissue-of-origin deconvolution15.
Three different readouts — fragment length, end motif, terminal methylation — all partly determined by one enzyme step chosen for reasons that had nothing to do with any of them.
The analysis is not a neutral observer either
The same systematic study found that trimming and alignment settings affected fragment length profiles more than the choice of library kit did11. That is not a rounding error in a preprocessing step. It means the most-cited fragmentomic feature is, in a measurable sense, downstream of a Nextflow parameter.
The root cause is a definitional ambiguity. In paired-end sequencing of short molecules, reads frequently read through into the adapter, and different aligners and toolkits define a “fragment” differently — outermost boundaries of the read pair, or the region between the start of the forward read and the end of the reverse. Choosing the wrong one produces artefacts that look like biology: peaks around 140 bp in three of the nine kits vanished once fragment length was calculated correctly, and a spurious paucity of reads below 150 bp appeared in three others11. A 140 bp shoulder in a length histogram is exactly the kind of feature a reader would interpret as sub-nucleosomal biology.
One routine habit deserves singling out. Filtering read pairs on the SAM “proper pair” flag is standard practice in genome sequencing. That flag is assigned by the aligner from an assumed approximately normal fragment-length distribution. Plasma cfDNA is not normally distributed — the di- and tri-nucleosomal peaks are the whole point — so applying the filter preferentially discards fragments in the di-nucleosomal range11. A quality-control step inherited from tissue sequencing silently deletes a feature that other groups are publishing as a biomarker.
| Feature | Most robust to | Dominated by | Practical implication |
|---|---|---|---|
| Global fragment size | Tube type, processing delay, centrifugation protocol | Fragment-length definition in code; library chemistry; single- vs double-stranded prep | Safe to compare across collection sites; not safe to compare across pipelines |
| Genome-wide fragmentation profile | Physiological variables tested to date | Collection tube and collection centre | Requires balanced collection across study arms |
| End motifs | Processing delay; centrifugation protocol | Library kit; end repair; collection centre | Cross-protocol comparison needs harmonisation or shared chemistry |
| Mitochondrial cfDNA | — | Centrifugation protocol; library kit | Report the mitochondrial read fraction as a QC metric |
Compiled from the studies cited below. “Robust to” means no significant difference was reported in the experiments described, at those sample sizes; it is not a claim that no effect exists at the scale a classifier can exploit.
Why correction is harder here than in expression data
The batch-effect toolkit exists and it works to a degree. Applying empirical Bayes harmonisation across published cfDNA datasets attenuated the clustering while preserving cancer signal11. Two caveats came with that result, and both matter.
The first is that batch effects across published studies could not be removed simply by reprocessing everything through one standardised pipeline11. Uniform analysis does not undo non-uniform wet-lab work. The second is structural: study, extraction kit and library kit were mutually confounded in the community data, so the batch variable is not cleanly separable into causes11. As the earlier piece on batch confounding in public data argued, when the technical variable and the biological variable occupy the same statistical space, no post-hoc procedure separates them; here that condition arises naturally, because a study is usually one protocol.
What makes fragmentomics harder than expression data is the size of the target. The authors note explicitly that in early-stage cancers with low tumour fraction, the differences from healthy controls are subtle, and whether harmonisation helps depends on the study design11. A correction that comfortably preserves a high-tumour-fraction signal can be the same magnitude as the signal you are trying to detect at screening-relevant tumour fractions. The regime where the assay would be most valuable is the regime where the correction is least safe.
What machine learning does with this
Fragmentomic classifiers take hundreds to thousands of correlated features and learn whatever separates the classes. A recent review of the field puts the concern directly: these confounding effects are especially salient when machine learning is involved downstream, because models amplify such biases, and their predictions become invalid when the confounders are distributed differently in deployment than in training16.
Readers of the recent piece on antimicrobial resistance prediction, and of the earlier one on population stratification, will recognise the shape. A model learns lineage instead of mechanism; a GWAS learns ancestry instead of association; a fragmentomics classifier learns the collection centre instead of the tumour. In every case the model is behaving correctly. It was given a variable that predicts the label, and it used it.
What to do about it
- Balance collection across study arms, not just demographics. If cases and controls come from different centres, tubes or time windows, the design cannot distinguish disease from procedure — and no analysis added later fixes that. Where balance is impossible, say so as a limitation rather than as a footnote in the methods.
- Treat pre-analytical metadata as variables. Tube type, time from draw to spin, centrifugation protocol, extraction kit, library kit, sequencer, pipeline version and reference build belong in the sample table with the clinical covariates, not in a laboratory logbook.
- Run the adversarial check before the diagnostic one. Train a classifier to predict collection centre, tube or library kit from the same feature matrix. If it succeeds, that information is available to the disease classifier too, and the reported AUC has an unknown component of it.
- Define the fragment explicitly, and do not inherit tissue defaults. Fix and document how fragment length is computed from paired reads, use library-specific adapter trimming, and do not filter on the “proper pair” flag — it assumes a fragment-length distribution that plasma cfDNA does not have.
- Match the library chemistry to the claim. If native 3′ ends, ultrashort fragments or terminal methylation are part of the argument, a double-stranded end-repaired protocol cannot support it. Single-stranded preparation is the instrument for those questions.
- Report the mitochondrial read fraction alongside coverage. It varies severalfold with library kit and moves with centrifugation, and it feeds genome-wide profiles.
- Tier the features by robustness in the report. A global size metric and an end-motif frequency do not deserve the same confidence when comparing across protocols. State which is which.
- Validate across protocols, not only across samples. The informative held-out set is a different kit, centre or pipeline. A random split of a single-protocol cohort measures reproducibility within a laboratory and says nothing about transfer.
- If harmonising, report what it did to the signal at low tumour fraction — not only that the overall separation survived.
The shape of the error
This series keeps arriving at the same structure from different directions. An FFPE artefact is real chemistry that happened in a cassette rather than in a patient. An ambient transcript is a real molecule counted against the wrong cell. A clonal haematopoiesis variant is genuinely somatic and from the wrong tissue. In each case the instrument worked and the inference did not.
Fragmentomics is the version where the artefact and the signal are not merely adjacent but identical in kind. A tumour shortens fragments; so does a size-biased bead chemistry. A nuclease leaves a characteristic terminal base; so does an end-repair enzyme. A dying cell type changes the ratio of short to long fragments in a genomic bin; so does the centre that drew the blood. There is no substitution class to filter, no provenance to check, no orthogonal assay that separates them, because the technical effect is not contaminating the measurement — it is the same measurement.
Which leaves study design carrying more weight than usual. The fragment profile is a faithful record of what happened to that blood: in the patient, in the tube, on the bench, and in the code. Deciding which of those the model is allowed to see is a decision made when the samples are collected, and not afterwards.
References
- Cristiano S, Leal A, Phallen J, et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature 2019;570:385–389. nature.com/articles/s41586-019-1272-6
- Mazzone PJ, Bach PB, Carey J, et al. Clinical validation of a cell-free DNA fragmentome assay for augmentation of lung cancer early detection. Cancer Discovery 2024;14(11):2224–2242. aacrjournals.org — Cancer Discov 14(11):2224
- Mouliere F, Chandrananda D, Piskorz AM, et al. Enhanced detection of circulating tumor DNA by fragment size analysis. Science Translational Medicine 2018;10:eaat4921. pubmed.ncbi.nlm.nih.gov/30404863
- Jiang P, Sun K, Peng W, et al. Plasma DNA end-motif profiling as a fragmentomic marker in cancer, pregnancy, and transplantation. Cancer Discovery 2020;10:664–673. pubmed.ncbi.nlm.nih.gov/32111602
- Han DSC, Ni M, Chan RWY, et al. The biology of cell-free DNA fragmentation and the roles of DNASE1, DNASE1L3, and DFFB. American Journal of Human Genetics 2020;106:202–214. pubmed.ncbi.nlm.nih.gov/31982378
- van der Pol Y, Moldovan N, Verkuijlen S, et al. The effect of preanalytical and physiological variables on cell-free DNA fragmentation. Clinical Chemistry 2022;68(6):803–813. academic.oup.com/clinchem/article/68/6/803
- Preprint version of ref. 6, from which the centre-clustering and t-SNE detail is drawn. bioRxiv 2021.09.17.460828. biorxiv.org — 2021.09.17.460828
- Hu X, Zhang H, Wang Y, et al. Effects of blood-processing protocols on cell-free DNA fragmentomics in plasma: comparisons of one- and two-step centrifugations. Clinica Chimica Acta 2024;560:119729. pubmed.ncbi.nlm.nih.gov/38754575
- Effects of collection and processing procedures on plasma circulating cell-free DNA from cancer patients. Journal of Molecular Diagnostics 2018. ncbi.nlm.nih.gov/pmc/articles/PMC6197164
- Evaluation of pre-analytical factors affecting plasma DNA analysis. Scientific Reports 2018;8:7375. nature.com/articles/s41598-018-25810-0
- Wang H, Mennea PD, Chan YKE, et al. A standardized framework for robust fragmentomic feature extraction from cell-free DNA sequencing data. Genome Biology 2025;26:141. genomebiology.biomedcentral.com — s13059-025-03607-5
- van der Pol Y, Moldovan N, Ramaker J, et al. The landscape of cell-free mitochondrial DNA in liquid biopsy for cancer detection. Genome Biology 2023;24:229. pubmed.ncbi.nlm.nih.gov/37828498
- Holistic determination of ends of cfDNA molecules. Cell Genomics 2026. cell.com/cell-genomics — S2666-979X(26)00004-2
- A review on the impact of single-stranded library preparation on plasma cell-free diversity for cancer detection. Frontiers in Oncology 2024;14:1332004. frontiersin.org — fonc.2024.1332004
- End-repair causes methylation underestimation in cell-free DNA sequencing libraries. 2026. sciencedirect.com — S2950195426000019
- Cell-free DNA fragmentomics in cancer. Cancer Cell 2025. cell.com/cancer-cell — S1535-6108(25)00398-8

