Copy Number Calling from Targeted Sequencing Panels: Bias, Baselines, and the Limits of Read Depth
Zetobit · Bioinformatics Insight Series
Copy Number Calling from Targeted Sequencing Panels: Bias, Baselines, and the Limits of Read Depth
Targeted panels were built to count variants, not chromosomes. When you ask them to call copy number, the absence of a genome-wide baseline turns ordinary noise into confident, wrong answers.
A single-exon deletion call lands in a clinical report. The evidence looks clean: mean coverage across the exon drops to roughly half of the surrounding capture, exactly what a heterozygous loss should produce. The variant scientist signs it out. Six weeks later, an orthogonal MLPA assay comes back flatly normal. Nothing was deleted. The panel had measured a dip in capture efficiency and called it biology.
Copy number variant (CNV) calling from targeted sequencing panels is now routine in clinical labs, and for good reason — panels are cheap, deep, and already validated for single-nucleotide variants. The temptation is to treat the same reads as a free CNV assay. But a panel is a fundamentally impoverished substrate for copy number, and most of its failure modes are silent: they produce plausible calls, not obvious errors.
Why depth is not copy number
The core assumption behind read-depth CNV calling is simple: sequencing coverage over a region is proportional to the number of copies of that region in the genome. Double the copies, double the reads. In whole-genome sequencing this holds reasonably well because the baseline — coverage everywhere else — is dense and continuous, so local deviations stand out against a stable genome-wide expectation.
A targeted panel destroys that baseline. You have coverage only where probes were placed, which might be a few hundred exons scattered across the genome. There is no continuous signal between targets, no smooth expectation to normalize against. Each target's depth has to be compared not to its neighbors on the chromosome but to the same target in a panel of reference samples. The quality of your CNV calls is therefore only as good as the stability of capture at each individual probe — and capture is not stable.
A panel doesn't sample the genome. It samples a few hundred places the chemistry happened to like that day.
The three biases that masquerade as deletions
GC content. Capture and amplification efficiency both track GC content. A GC-rich exon captures less efficiently and reads shallower; a GC-poor one reads deeper. If your reference panel and your test sample were prepared in different batches — different reagent lots, different operators, different days — the GC response curve shifts, and every high-GC exon can appear to lose a copy at once. Without explicit GC correction1, this is one of the most common sources of false single-exon calls.
Probe-level capture variability. Individual probes have idiosyncratic, reproducible efficiencies driven by their sequence, secondary structure, and proximity to repeats. A probe that consistently under-captures looks like a recurrent heterozygous deletion across many samples. Labs that don't build a probe-specific reference — pooling many normals to learn each target's expected behavior2 — inherit these artifacts as calls.
Batch and reference mismatch. CNV callers for panels (CNVkit, DECoN, ExomeDepth, and similar tools2,3) all depend on a reference set of normal samples processed the same way as the test. When the test sample's library prep diverges from the reference pool, the normalization breaks systematically, not randomly — producing coherent, deletion-shaped signals across correlated targets.
The single-exon problem
The smaller the event, the worse the substrate. A deletion spanning many contiguous targets accumulates evidence across all of them — a coherent shift across ten exons is hard to fake. A single-exon deletion has exactly one data point, and that point competes directly with probe-specific noise. This is why most panel CNV pipelines quietly restrict confident reporting to multi-exon events and treat single-target calls as low-confidence flags requiring orthogonal confirmation4. A pipeline that reports single-exon CNVs at the same confidence as multi-exon ones is not being sensitive; it is failing to model its own noise floor.
What actually works
None of this means panels can't yield useful CNV information — it means the information has to be earned. In practice that means: build a large, batch-matched reference of normals processed identically to test samples; apply explicit GC and probe-level bias correction rather than assuming raw depth is proportional to copy number; hold single-exon calls to a higher evidence bar than multi-exon events; and confirm clinically actionable losses with an orthogonal method — MLPA, ddPCR, or array — before sign-out. The panel is a screening instrument for copy number, not a confirmatory one.
The deletion that started this piece was never real. What was real was a batch difference between the test library and the reference pool, concentrated in the GC-rich exons the chemistry handled worst. The depth didn't lie so much as it answered a question it was never designed to be asked.
— Zetobit
References
- Benjamini Y, Speed TP. Summarizing and correcting the GC content bias in high-throughput sequencing. Nucleic Acids Research. 2012;40(10):e72. doi:10.1093/nar/gks001
- Talevich E, Shain AH, Botton T, Bastian BC. CNVkit: Genome-wide copy number detection and visualization from targeted DNA sequencing. PLOS Computational Biology. 2016;12(4):e1004873. doi:10.1371/journal.pcbi.1004873
- Fowler A, Mahamdallie S, Ruark E, et al. Accurate clinical detection of exon copy number variants in a targeted NGS panel using DECoN. Wellcome Open Research. 2016;1:20. doi:10.12688/wellcomeopenres.10069.1
- Roca I, González-Castro L, Fernández H, Couce ML, Fernández-Marmiesse A. Free-access copy-number variant detection tools for targeted next-generation sequencing data. Mutation Research/Reviews in Mutation Research. 2019;779:114-125. doi:10.1016/j.mrrev.2019.02.005

