What the QC Report Omits
What the QC Report Omits
Why “passed” describes the run, not your gene.
The data comes back, and at the front of it sits a page nobody plans to spend time on. A short column of numbers, a limit printed beside each one, and a word at the bottom: passed. Two seconds of attention is about the right allocation, provided those two seconds read the page as what it is. The difficulty is that “passed” tends to be heard as a broad reassurance about the whole result, and what it actually certifies is far narrower than that.
A pass is a comparison, not a measurement. Every value on the page was checked against a threshold, and the thresholds belong to the laboratory that issued the report. The ACMG technical standard requires laboratories to develop quality-control metrics from performance parameters established during their own assay development, and to apply those metrics in validation and in routine production alike.1 The AMP/CAP standard for sequencing pipelines adds the part that matters most to a reader: quality metrics should be evaluated against the sensitivity, specificity and predictive values the test is intended to achieve.2 Put those together and a pass carries a precise, bounded claim — this run fell inside the range in which this laboratory characterized this assay for this purpose. It is not a score, and it does not travel between laboratories. A lower threshold at one lab is not a lower standard; it may be a different assay answering a different question. The bars are set by what the assay has to be able to detect: a germline panel reporting heterozygous variants and a somatic panel expected to find a mutation present in a twentieth of the cells do not need the same depth, and the difference between their thresholds reflects that, not a difference in rigor.
Research deliveries invert the problem. The summary that arrives with data from a sequencing core often carries no thresholds at all — the green and amber marks are a tool's generic defaults, not a decision anyone made about this experiment, this tissue or this question. That version of the page is worth reading as a description rather than a verdict, and its cheerfulness carries no information about whether the data can support the comparison you intend to make with it.
What follows are four things the page does not say, roughly in the order they tend to matter.
The distribution behind the average. Mean depth and percent-of-target-above-threshold are aggregates over hundreds of thousands of positions, and an aggregate is defined entirely by what it aggregates over. When one group compared thirty-six clinical exomes from three laboratories, the headline figures for two of them were effectively identical: 96.49% and 96.54% of coding nucleotides at 20× or better, against 91.68% for the third. Underneath those numbers, the count of genes covered completely across their coding sequence was 12,184, 11,687 and 5,989.3 The same summary statistic, sitting on top of very different sets of genes. Variability within a single laboratory, sample to sample, ran to coefficients of variation of 29%, 13% and 37% — one assay, one lab, different patients. And the spread widened as the gene set narrowed: for a subset of epilepsy genes and for the secondary-findings genes, the coefficient of variation at one laboratory reached 60% and 71%.3 The exome-wide average is least informative exactly where it is most often used, which is when the reader cares about a short list of genes. This is why the AMP/CAP standard asks laboratories to report the distribution of quality metrics across the regions analyzed, not only their central values.2
Where the shortfall landed. If 96% of the target met the bar, the page has told you nothing about the other 4% except its size — and size is the least useful thing about it. Location is what matters, and location is not stable from run to run. A 2026 analysis of exome and genome data found low-coverage regions to be largely batch-specific: consistency within a sequencing batch was high, but only a small subset of gaps was shared across four independent runs using the same capture kit.4 Those gaps fell in disease-associated genes and occasionally on positions of known pathogenic variants. Two consequences follow for anyone holding a report. A coverage summary from a previous run does not describe this sample. And a gene that was fully covered for the last patient can be a hole for this one. Coverage is a property of the run, not a property of the assay.
The checks that were never run. The third omission is the one the page cannot signal, because it takes the form of a silence. A quality-control summary lists the results of checks that were performed. ACMG's standard notes that QC metrics should be chosen to monitor sample integrity as well as data integrity1 — two different questions that a single green word tends to merge. Data metrics describe the sequence: how much of it there is, how confident the instrument was in each base, how much of the library was redundant. Every one of them can be excellent on a specimen that belongs to someone else, or that carries a second person's DNA alongside the patient's. If no line for identity concordance or contamination estimate appears on the page, the page is not telling you those were clean. It is not telling you anything about them at all. The same silence covers the run's history: whether a control was processed alongside this specimen, what became of the other samples that shared the flow cell, and whether this was a first attempt or a repeat after something failed are all part of the laboratory's record, and none of them are on the summary.
That depth is not sensitivity. The fourth is a conversion readers perform without being asked: depth becomes confidence. It does not hold as tightly as the arithmetic suggests. When three vendors sequenced the same trios, mean depths came out at 189×, 125× and 38×; across nine individuals, 55 of the 56 genes then recommended for secondary-findings reporting were completely covered at 10× by every vendor — but only 56 of 63 pharmacogenes were, with substantial variation between individuals.5 The authors' recommendation was that reports should estimate sensitivity directly rather than leave the reader to infer it from depth,5 which is a courteous way of observing that depth was already being read as an answer to a question it does not answer. A high mean depth makes good coverage likely everywhere. It does not make it a fact anywhere in particular.
| Metric | What it establishes | What it does not establish |
|---|---|---|
| Mean target depth | Average reads per position across the whole target | That any particular position was adequately covered |
| % target ≥ threshold | The fraction of the target that cleared the bar | Which regions made up the fraction that did not |
| Bases ≥ Q30 | The instrument's confidence in the base calls | That reads were placed correctly, or came from this patient |
| Duplicate rate | How much of the library was redundant overall | The independent coverage available at a given locus |
| Overall: PASS | That no monitored metric crossed a laboratory threshold | That nothing unmonitored went wrong |
None of this makes the page useless. It makes it a filter rather than a certificate. It is a truthful, well-defined statement that a run cleared the boundaries a laboratory drew around its own assay, and it catches gross failures reliably — which is exactly what you want established before anyone spends time on what comes next. It was simply never designed to answer the question most readers bring to it, which is whether the particular result in front of them could have been missed.
One request closes most of that gap, and it is usually easy to satisfy: ask for the per-gene or per-exon coverage for the genes relevant to your question, in this sample. The information exists — laboratories are expected to characterize the capture efficiency of their method from their own validation studies,6 and for exome and genome tests ACMG states plainly that complete coverage is not expected in the first place.1 Ask early rather than late. The question costs a few minutes while a sample is still moving through the laboratory, and considerably more once a negative result is already sitting in a chart or a manuscript.
Three questions to ask
- What thresholds define a pass for this assay, and what level of sensitivity were they set to deliver?
- Which regions of the genes I care about fell below threshold in this sample — not on average, and not during validation?
- Were specimen identity and contamination assessed for this run, and if so, where do those results appear?
References
- Rehder C, Bean LJH, Bick D, et al. Next-generation sequencing for constitutional variants in the clinical laboratory, 2021 revision: a technical standard of the American College of Medical Genetics and Genomics (ACMG). Genet Med. 2021;23(8):1399–1415. doi:10.1038/s41436-021-01139-4
- Roy S, Coldren C, Karunamurthy A, et al. Standards and Guidelines for Validating Next-Generation Sequencing Bioinformatics Pipelines: A Joint Recommendation of the Association for Molecular Pathology and the College of American Pathologists. J Mol Diagn. 2018;20(1):4–27. doi:10.1016/j.jmoldx.2017.11.003
- Gotway G, Crossley E, Kozlitina J, et al. Clinical Exome Studies Have Inconsistent Coverage. Clin Chem. 2020;66(1):199–206. doi:10.1093/clinchem.2019.306795
- Iovino E, De Masi C, Ballestrazzi A, et al. Distribution of Sequencing Coverage Gaps in Exomes and Genomes: Potential Implications for Diagnostic Accuracy in Neurodevelopmental Disorder Genes. Genes. 2026;17(3):269. doi:10.3390/genes17030269
- Kong SW, Lee I-H, Liu X, Hirschhorn JN, Mandl KD. Measuring coverage and accuracy of whole-exome sequencing in clinical context. Genet Med. 2018;20(12):1617–1626. doi:10.1038/gim.2018.51
- Rehm HL, Bale SJ, Bayrak-Toydemir P, et al. ACMG clinical laboratory standards for next-generation sequencing. Genet Med. 2013;15(9):733–747. doi:10.1038/gim.2013.92

