What Counts as the Same Event: Validating Copy Number and Structural Variant Detection

What Counts as the Same Event: Validating Copy Number and Structural Variant Detection — Zetobit
Zetobit The Validation File
CLAIMS EVIDENCE UPKEEP THE VALIDATION FILE Validating CNVs and SVs One call makes three claims at once. A validation measures whichever one the scoring rule scores. Kanna Nandakumar, PhD ZETOBIT

The Validation File

What Counts as the Same Event: Validating Copy Number and Structural Variant Detection

One CNV call makes three claims at once. A validation measures whichever one the concordance rule happens to score — which is why the rule has to be written down before any samples run.

A laboratory finishes validating its panel for sequence variants. Sensitivity is 99.6 percent, the concordance table is clean, the file closes. Then the same panel starts reporting deletions and duplications, and the template that produced that clean table stops working.

Not because copy number calling is harder, though it is. Because the question the table answers no longer has an unambiguous form. For a sequence variant, did the pipeline detect it? compares a position and a base: two methods agree or they do not. A copy number call is a region, a direction and a dosage. Two methods can both be substantially right about the same event and still return different coordinates, different sizes, and — under a strict comparison — no match at all.

That is not a nuisance to be handled at the analysis stage. It is a design decision that belongs in the validation plan, because the definition chosen there determines the number that later appears in the test description and in the file an inspector reads.

The obligation is parallel, not inherited

A laboratory running a test of its own design, or a cleared test it has modified, must establish performance specifications for the test as it actually performs it.1 The AMP/CAP standard for bioinformatics pipeline validation states the scope plainly: validation must be appropriate for the intended clinical use, the specimen, and the variant types the test detects.2 ACMG's 2023 statement on germline structural variant detection converts that into an instruction — the validation should encompass every type of SV the assay targets, the size range should be defined for each type, and a variant type added later must be validated when it is added.3

Taken together, these say something laboratories often discover late: CNV and SV detection is not a subsection of the sequence-variant validation but a parallel one, with its own samples, concordance definition, quality metrics and thresholds. ACMG is explicit on the last point — the cutoffs for SV analysis are separate from those applied to sequence variants, and the laboratory needs a written procedure for the sample whose sequence QC passes and SV QC does not.3

One call, three claims

A single figure cannot carry the weight because a copy number call asserts three things at once: that an event is present, that it spans approximately these coordinates, and that the copy state is this. Different evidence supports each. Depth-of-coverage analysis compares binned read counts against a reference and is what detects aneuploidies, large events and repeat-mediated CNVs — but it places boundaries only to the resolution of a bin. Split reads and discordant pairs locate breakpoints, and make balanced rearrangements and small events visible at all — but they do not measure dosage. ACMG recommends combining the classes precisely because each is blind to what the other sees.3

For the file, the consequence is concrete. A sensitivity figure for CNV detection is a claim about one of those three assertions, and the validation plan has to say which.

One CNV / SV call CLAIM 1 Something is here CLAIM 2 It spans these edges CLAIM 3 The copy state is N EVIDENCE Read depth across consecutive bins Split reads and discordant pairs Sample-to-reference ratio; VAF in region COMPARATOR CMA, MLPA, qPCR Breakpoint PCR + Sanger; dense array CMA, qPCR, karyotype / FISH SCORED BY A RULE YOU CHOOSE any overlap at all? reciprocal overlap ≥ x%? direction only, or level? WHAT THE VALIDATION HAS TO PRODUCE concordance definition per comparator pair · stratified resolution statement challenging-region catalog · confirmation policy · recorded exclusions
A copy number call is three assertions travelling together. Each is supported by a different kind of evidence and adjudicated by a different comparator, and each is scored by a rule the laboratory selects. The rule is not a downstream detail — it is what turns three assertions into one number, and it belongs in the validation plan rather than in the analyst's judgement.

The scoring rule decides the number

The clearest recent demonstration comes from a 2025 benchmark of short-read CNV callers built explicitly around clinical needs. Evaluated genome-wide against the Genome in a Bottle HG002 truth set, the most sensitive configuration reached 83% sensitivity at 30% precision. Evaluated on the same kind of data against 25 characterized cell lines across a 184-gene panel, the same tool reached 100% sensitivity.4

Much of that gap is definitional. The authors state that a call was scored correct if it overlapped at least one base pair of a coding exon and matched the dosage direction, deliberately declining to assess breakpoint accuracy or length overlap on the grounds that clinical reporting turns on whether coding sequence is disrupted, not on where exactly the edges fall.4 Both numbers are defensible. They answer different questions, and a reader cannot tell from the number which question was asked.

One case from that study is worth carrying into any validation plan. A real deletion in ACTN2 was scored a false positive because its right breakpoint was called too far out, pushing the event across an exon boundary it should not have crossed.4 Presence: correct. Copy state: correct. Boundary: wrong. Under that scoring rule, the whole call counts against precision — and under a looser rule, it counts for sensitivity.

This is why ACMG asks laboratories to define their concordance metrics in the validation plan rather than settle them afterwards, and supplies worked values to start from: a reciprocal overlap of roughly 75% when comparing CNVs between chromosomal microarray and genome sequencing, 90% when comparing genome with genome, and — when comparing genome with exome — a definition based on the affected genes and exons rather than on genomic coordinates at all. The same guidance suggests aiming the concordance target, above 95%, at clinically significant CNVs rather than at every event the assay emits.3

The comparator is not the truth

The other half of the problem is what you compare against. Microarray does not see balanced rearrangements. MLPA sees only where its probes sit. Karyotype is silent below what it can resolve. Each comparator answers a subset of the three claims, and a validation that treats one as ground truth inherits its blind spots without recording them.

Even reference materials are narrower than they appear. The Genome in a Bottle SV benchmark contains 12,745 isolated, sequence-resolved insertions and deletions of at least 50 bp, with Tier 1 regions covering 2.51 Gbp of the genome.5 It was built by filtering out clustered and complex calls so that scoring would be unambiguous — which is exactly what makes it usable, and also what makes it thinnest in the contexts where clinical copy number work is hardest. The 2025 benchmark above could not assess performance in challenging medically relevant genes at all, for want of samples carrying pathogenic events there.4

How many positives a validation needs is a question of its own, and one this series takes up separately. Composition matters at least as much as count: the recommendation is for samples spanning a range of sizes, copy number states and genomic contexts — segmental-duplication-mediated versus unique breakpoints, high and low GC content — including the recurrent microdeletion and duplication regions the assay claims to cover.3,8 Where physical samples do not exist, in silico data fills part of the gap. The AMP/API/CAP working group is clear about the limit: simulated data generally overestimates performance because it cannot reproduce all the biases and errors present in real data, and performance against a simulated set should not be conflated with analytical sensitivity or specificity.6 It supplements physical samples; it does not stand in for them.

And when two methods disagree, neither one is the referee. The AMP/NSGC confirmation recommendations hold that a discrepancy should be investigated rather than resolved by defaulting to whichever assay is treated as primary, and that the confirmation policy itself, exceptions included, belongs in the report.7

Half of what you validated is data

For any target-enrichment assay, the copy number call depends on a reference set — a pool of samples the test sample's coverage is measured against. That pool is part of the assay. ACMG asks that it be defined during validation, built from the same specimen type and processed like clinical samples, kept free of aneuploidies, and governed by a documented rule for when and how it is rebuilt: a static reference set is sensitive to any change in extraction, library preparation or pipeline processing, and there is no established guidance on how long one stays valid.3

This is the part that most often sits outside a laboratory's change-control paperwork. A container digest pins the software; it does not pin the reference pool. What was validated is code plus a dataset, and only one of the two is version-controlled by default. What counts as a change large enough to trigger revalidation is a topic of its own, later in this series — but the reference set has to be named in the file before that question can even be asked.

Resolution is a statement, not a number

Finally, the part clinicians actually read. ACMG asks that resolution be defined during validation and stated in the test description, expressed as a minimum number of affected exons, baits or bins for targeted assays or as a size for genome sequencing; that it be defined separately for gains and losses, since duplications are generally harder to detect than deletions; and it accepts that coverage variability may make a single precise figure impossible across an exome, so that a laboratory may legitimately quote higher resolution in high-coverage unique regions and lower resolution elsewhere.3

That is not hedging. It is the honest form of the claim, and it has a documentary consequence. The validation has to produce a catalog of challenging regions and recurrent technical artifacts, a recorded decision on whether each is masked before analysis or flagged during it, and disclosure of the regions where copy number cannot be determined at all — particularly those containing dosage-sensitive genes.3

A sequence-variant validation can end in a number. A copy number validation ends in a set of definitions, a stratified account of what the assay resolves where, and an explicit list of what was not assessed. That is a longer document. It is also the only version an inspector, or a clinician reading a limitations paragraph two years from now, can act on.

What goes in the file

  • A CNV/SV validation plan naming each variant type the assay targets and the size range claimed for each, dated before the validation samples were run.
  • The concordance definition: the reciprocal overlap threshold, or the gene- and exon-level rule, applied to each comparator pair — plus the target set for clinically significant events.
  • A sample manifest recording each event's size, copy state and genomic context, the recurrent microdeletion and duplication regions covered, and what each comparator method could and could not have seen.
  • Any in silico data used, with what it stood in for and why physical samples were unavailable.
  • The reference set specification: composition, specimen type, static or dynamic, and the documented rule for rebuilding it.
  • The resolution statement in the form that goes into the test description — separately for gains and losses, stratified by region where a single figure would misrepresent the assay.
  • The catalog of challenging regions and recurrent artifacts, with the masked-or-flagged decision recorded for each.
  • SV-specific QC metrics and cutoffs, and the procedure for samples that pass sequence QC but fail SV QC.
  • The orthogonal confirmation policy: which findings it covers, which are reported on primary evidence, and what happens when the two methods disagree.
  • Exclusions — the variant types, regions and genes not assessed during validation, and the reason each was left out.

References

  1. 42 CFR § 493.1253 — Standard: Establishment and verification of performance specifications. ecfr.gov
  2. Roy S, Coldren C, Karunamurthy A, et al. Standards and Guidelines for Validating Next-Generation Sequencing Bioinformatics Pipelines: A Joint Recommendation of the Association for Molecular Pathology and the College of American Pathologists. J Mol Diagn. 2018;20(1):4–27. doi:10.1016/j.jmoldx.2017.11.003
  3. Raca G, Astbury C, Behlmann A, et al.; ACMG Laboratory Quality Assurance Committee. Points to consider in the detection of germline structural variants using next-generation sequencing: A statement of the American College of Medical Genetics and Genomics (ACMG). Genet Med. 2023;25(2):100316. doi:10.1016/j.gim.2022.09.017
  4. De La Vega FM, Irvine SA, Anur P, et al. Benchmarking of germline copy number variant callers from whole genome sequencing data for clinical applications. Bioinform Adv. 2025;5(1):vbaf071. doi:10.1093/bioadv/vbaf071
  5. Zook JM, Hansen NF, Olson ND, et al. A robust benchmark for detection of germline large deletions and insertions. Nat Biotechnol. 2020;38(11):1347–1355. doi:10.1038/s41587-020-0538-8
  6. Duncavage EJ, Coleman JF, de Baca ME, et al. Recommendations for the Use of in Silico Approaches for Next-Generation Sequencing Bioinformatic Pipeline Validation: A Joint Report of the Association for Molecular Pathology, Association for Pathology Informatics, and College of American Pathologists. J Mol Diagn. 2023;25(1):3–16. doi:10.1016/j.jmoldx.2022.09.007
  7. Crooks KR, Farwell Hagman KD, Mandelker D, et al. Recommendations for Next-Generation Sequencing Germline Variant Confirmation: A Joint Report of the Association for Molecular Pathology and National Society of Genetic Counselors. J Mol Diagn. 2023;25(7):411–427. doi:10.1016/j.jmoldx.2023.03.012
  8. Marshall CR, Chowdhury S, Taft RJ, et al.; Medical Genome Initiative. Best practices for the analytical validation of clinical whole-genome sequencing intended for the diagnosis of germline disease. NPJ Genom Med. 2020;5:47. doi:10.1038/s41525-020-00154-9
Previous
Previous

Where the Wet Lab Ends and Pipeline Validation Begins

Next
Next

How Many Positives Do You Need: Why 25, 59 and 600 Are All Correct Answers