Tumor Purity and Ploidy: Why the Number Every Somatic Result Depends On Has More Than One Valid Answer

Tumor Purity and Ploidy: Why the Number Every Somatic Result Depends On Has More Than One Valid Answer — Zetobit
Zetobit.
BIOINFORMATICS INSIGHT SERIES Tumor Purity and Ploidy Why the number every somatic result depends on has more than one valid answer ONE OBSERVED COVERAGE PROFILE SOLUTION A purity 0.72 ploidy 2.0 near-diploid · no WGD SOLUTION B purity 0.41 ploidy 3.9 genome-doubled Both fit the data. One gets reported. ZETOBIT. Kanna Nandakumar, PhD

Bioinformatics Insight Series

Tumor Purity and Ploidy: Why the Number Every Somatic Result Depends On Has More Than One Valid Answer

Purity and ploidy are not measured. They are fit — jointly, from the same coverage data, in a model where different combinations can explain the observations equally well. The tool reports one solution. The data often supported several.

A tumor sequencing report is full of numbers that look like observations. Copy number 4 at ERBB2. This mutation is clonal, that one subclonal. Tumor mutational burden of 11 per megabase. Loss of heterozygosity at BRCA1.

Almost none of those are read directly off the data. They are computed downstream of two parameters — the fraction of cells in the specimen that are tumor, and the average copy number of the tumor genome — and those two parameters are themselves estimates, produced by fitting a model to observed coverage and allele frequencies.

The awkward part is not that they're estimated. It's that they are estimated together, from the same evidence, in a system where trading one against the other can produce nearly the same observed data. A tumor at 72% purity with a normal diploid genome and a tumor at 41% purity that has doubled its entire genome will look remarkably similar in a coverage profile. The algorithm must choose. Its choice determines the copy number of every segment, the clonality of every mutation, and whether the report says whole-genome doubling occurred.

The identifiability problem, by its own name

This is not a subtle or contested point in the methods literature — it has a name and appears in the motivation section of nearly every tool paper in the field. The identifiability problem refers to the fact that different combinations of tumor purity and ploidy often explain the sequencing data equally well.

The reason is structural. What a sequencer observes at a genomic segment is a depth ratio, and that ratio is a function of three unknowns simultaneously: the absolute copy number at that segment, the proportion of tumor cells contributing signal, and the overall ploidy that sets the baseline against which the ratio is scaled. One observation, three unknowns. The system is underdetermined — in the phrasing of one method paper, the deconvolution problem in general has multiple equivalent solutions.

Every tool resolves this the only way it can: by adding an assumption that is not in the data. Reviewing the field, one paper catalogues exactly this — methods based on B-allele frequencies at somatic mutations effectively assume tumor cell ploidy equals 2; CNAnorm prefers the solution closest to diploid; ABSOLUTE brings in karyotype data alongside coverage; and some approaches output all optimal solutions rather than choosing.

Those assumptions are reasonable. They are also invisible in the output. A report stating "purity 0.41, ploidy 3.9" carries no indication of whether that was the unambiguous answer or one of several the model ranked closely — nor which prior tipped the decision.

WHAT IS OBSERVED depth ratio at this segment = 1.38 · B-allele frequency = 0.34 FIT A purity 0.72 · ploidy 2.0 copy number here: 3 mutation CCF: 0.94 → clonal WGD: not called a gain on a diploid background FIT B purity 0.41 · ploidy 3.9 copy number here: 5 mutation CCF: 0.52 → subclonal WGD: called a relative loss on a doubled genome Same reads. Same evidence. Two internally consistent readings of the tumor. Illustrative values — the point is the structure, not these specific numbers.

Figure 1. A depth ratio does not determine copy number on its own; it constrains a combination of copy number, purity, and ploidy. Choosing a different point in that space changes the integer copy number assigned to a segment, the cancer cell fraction computed for mutations on it, and whether the genome is declared doubled. The two readings are not "one right and one wrong" from the data alone — they are two solutions the model must choose between using an assumption.

Whole-genome doubling is the axis where it bites

The doubling ambiguity is not a theoretical corner case. Whole-genome doubling was identified in the tumors of nearly 30% of 9,692 prospectively sequenced advanced cancer patients, varying by lineage and molecular subtype and arising early in carcinogenesis. In some settings it runs higher still: among non-hypermutated colorectal cancers, one study found WGD in 54% of early-onset versus 38% of late-onset disease.

So the model is routinely asked to distinguish two genuinely common states, and the distinction is exactly the one the mathematics blurs. This is compounded by how existing tools handle it. One method paper notes that existing approaches exclude WGD from model selection, or do not consider the trade-off between inferring multiple tumor clones versus inferring a WGD, or evaluate solutions using purity and ploidy — composite parameters that average over the unknown copy numbers and proportions, and so may not adequately distinguish between multiple equally plausible solutions.

That last observation is worth pausing on. Purity and ploidy are summaries. Two very different underlying genome configurations can produce the same summary pair, which means selecting a model by comparing those two numbers is a coarser test than the problem requires.

The same reads support both readings. What separates them is a prior, and the prior does not appear in the report.

How the error propagates

If purity were only used to annotate a report, a wrong value would be a cosmetic problem. It isn't. Purity is a divisor.

Cancer cell fraction — the proportion of tumor cells carrying a mutation, and the basis for calling anything clonal or subclonal — is not observed directly. It must be inferred from the variant allele fraction, the mutation multiplicity, the local copy number, and the tumor purity of the bulk sample. Every one of those inputs except VAF is itself an estimate, and two of them come from the purity–ploidy fit.

The direction of the resulting error is predictable, which is useful diagnostically. Underestimated purity and overestimated ploidy inflate CCF estimates; overestimated purity and underestimated ploidy compress the range of CCFs. A compressed CCF range makes a polyclonal tumor look clonal. An inflated one manufactures subclones.

The downstream evidence bears this out. Work on QC for copy number and CCF found that uncertainty metrics identified spurious subclonal clusters explained by miscalled CCFs, and framed this explicitly as the importance of QC to avoid propagating errors into downstream analyses. The same work found purity itself governs how much of the data is usable: in one cohort at around 50% purity, roughly 15% of mutations could not be assigned a reliable CCF, while in another cohort at similar purity around 40% could not — a difference the authors attribute to sequencing depth.

Benchmarking of subclonal deconvolution pipelines points the same direction: increasing either purity or purity-corrected sequencing depth improves accuracy, while tumour complexity itself does not drive accuracy. What limits subclonal inference is not how complicated the tumor is; it is how much tumor signal is present and how well the purity correction was done.

Table 1. Reported quantities that are computed downstream of the purity–ploidy fit rather than measured. An error in the fit moves all of them together, in a correlated way.
Reported quantity Dependence on the fit Effect of a wrong solution
Integer copy number Depth ratio scaled by purity and ploidy baseline Amplification/deletion thresholds cross or fail to cross
Cancer cell fraction VAF adjusted by local copy number, multiplicity, purity Clonal↔subclonal reclassification; phantom subclones
Whole-genome doubling Directly the ploidy branch of the solution A prognostically meaningful call flips
Loss of heterozygosity Allele-specific copy number from the same fit LOH-dependent eligibility calls change
Subclonal architecture Built on the CCF distribution Cluster count and evolutionary ordering shift

Methods disagree, and they disagree most where it matters

Because each tool encodes a different prior, they can return different answers on the same sample — and the disagreement concentrates in low-purity and high-aneuploidy samples, which are the clinically difficult ones.

A systematic assessment of purity estimation methods found that several failed to produce an estimate at all on a complete dataset, with missing rates including ASCAT at 4.8%, CLONET at 12.9%, and ABSOLUTE at 14.1%. The authors noted that the methods that failed were DNA-based, suggesting intrinsic limitations in estimating tumor purity from DNA-based assays in that setting. A method declining to answer is arguably the honest outcome; the risk is a pipeline that treats a missing estimate as a failure to be defaulted around.

Cross-method ploidy comparison shows the same pattern. In a benchmark across TCGA tumors, one tool was reported as quite good at predicting ploidy but weaker on purity — the two parameters are not equally recoverable, and a method can be trusted for one and not the other. The same work notes that where multiple valid purity–ploidy options exist, which it calls a known issue, its own approach defaults to the simplest solution. That is a defensible convention and it is still a convention.

What this asks of an implementation

Record purity and ploidy as pipeline parameters with provenance

Tool, version, and the specific solution selected — including the fit statistic and whether alternatives were close — belong in the record alongside the calls they produced. A downstream number derived from a purity of 0.41 means something different if the runner-up solution had purity 0.72.

Inspect the alternative solutions, don't just take the top fit

Tools that expose the solution space make this possible; some output all optimal solutions rather than one. Where a second solution sits close in likelihood, that sample warrants review rather than automated reporting — particularly if the two solutions differ on WGD.

Bring in evidence the fit does not have

The identifiability problem is resolved by adding information, which is precisely what the better methods do — integrating loss of heterozygosity alongside copy number, or incorporating karyotype data. Pathology cellularity estimates, orthogonal ploidy measurement, and multi-region or matched samples all serve the same function: constraining a system the coverage data alone leaves underdetermined.

Set purity floors per assay and honor them

Below some purity, CCF simply cannot be assigned reliably for a large share of mutations, and the fraction affected depends on depth. That threshold is a property of the assay and should gate subclonal reporting explicitly rather than being discovered case by case.

Treat clonality claims as the most fragile output

Copy number errors are bounded by integer states. CCF errors are continuous and directional, and they determine claims — clonal driver, subclonal resistance mutation, branching evolution — that read as biological conclusions rather than model outputs. They deserve the most conservative reporting.

The short version

Tumor purity and ploidy are jointly fit from data that does not uniquely determine them. Every tool breaks the tie with an assumption, and the assumption does not travel with the result. Because purity is a divisor for copy number and cancer cell fraction, a plausible-but-wrong solution produces a complete, internally consistent, and entirely different picture of the tumor.

An estimate wearing the clothes of a measurement

The recurring theme in this series is outputs that present as observations when they are inferences — a basecall that is a model's reading of a signal, a classification that is a dated claim about evidence. Purity and ploidy belong to that family, with a distinguishing feature: the ambiguity is not a limitation anyone is hiding. It has a name, it is stated in the introduction of the tool papers, and the field has built an entire methods literature around resolving it.

What gets lost is the transmission. By the time purity has become a copy number, and the copy number has become a cancer cell fraction, and the cancer cell fraction has become the sentence "this resistance mutation is subclonal," the fact that a model chose between competing solutions has disappeared from view. The number that carried the uncertainty is three steps upstream and is not in the report.

Keeping it in the report is the whole of the fix.

Zetobit builds and validates CAP/CLIA-compliant NGS pipelines, including somatic workflows where purity–ploidy solutions, alternative fits, and purity-dependent reporting thresholds are part of the validation record. If you are standing up somatic copy number or subclonal analysis, or auditing clonality calls across a cohort, we're happy to talk.

References

  1. Li Y, Xie X. Deconvolving tumor purity and ploidy by integrating copy number alterations and loss of heterozygosity. Bioinformatics 30(15):2121–2129 (2014). doi:10.1093/bioinformatics/btu174.
  2. Bielski CM, Zehir A, Penson AV, et al. Genome doubling shapes the evolution and prognosis of advanced cancers. Nature Genetics 50:1189–1195 (2018). doi:10.1038/s41588-018-0165-1.
  3. Zaccaria S, Raphael BJ. Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data (HATCHet). bioRxiv preprint; later Nature Communications 11:4301 (2020). doi:10.1038/s41467-020-17967-y.
  4. Antonello A, Bergamin R, Calonaci N, et al. Computational validation of clonal and subclonal copy number alterations from bulk tumor sequencing using CNAqc. Genome Biology 25:38 (2024). doi:10.1186/s13059-024-03170-5.
  5. Tanner G, Westhead DR, Droop A, Stead LF. Benchmarking pipelines for subclonal deconvolution of bulk tumour sequencing data. Nature Communications 12:6396 (2021). doi:10.1038/s41467-021-26698-7.
  6. Luo Z, Fan X, Su Y, Huang YS. Accurity: accurate tumor purity and ploidy inference from tumor-normal WGS data by jointly modelling somatic copy number alterations and heterozygous germline single-nucleotide-variants. Bioinformatics 34(12):2004–2011 (2018). doi:10.1093/bioinformatics/bty043.
  7. Systematic assessment of tumor purity and its clinical implications. JCO Precision Oncology (2020). doi:10.1200/PO.20.00016.
  8. Tumor ploidy determination in low-pass whole genome sequencing and allelic copy number visualization using the Constellation Plot (BACDAC). BMC Bioinformatics (2025). doi:10.1186/s12859-025-06139-8.
Previous
Previous

Immune Repertoire Sequencing: Why a Clone Frequency Is an Amplification Artifact and a Clone Is a Threshold You Chose

Next
Next

Transcript Abundance and Protein Abundance: What an RNA-Seq Result Does Not Establish