What “99.5% Sensitivity” Is a Claim About
What “99.5% Sensitivity” Is a Claim About
The quantity, the denominator, and the interval — three choices made before the number exists, none of them visible in it.
A validation study closes and a number appears in the summary: 99.5% sensitivity. Over the following year that number travels. It goes into a brochure, a scope of work, a payer packet, a data-room slide. It arrives at each stop intact and stripped of the things that made it true: the comparator it was measured against, the variants it was measured over, and the number of observations behind it. Somewhere around the third stop it stops being a claim and becomes an attribute of the assay, like a catalog number.
The previous piece in this series asked which obligation a laboratory is actually under: verify, establish, or demonstrate clinical validity. This one assumes that question is settled and asks a narrower one: when the establishment work produces a performance figure, what exactly has been established, and what has to be in the file for it to survive somebody asking?
Three choices are made before the number is computed. Each of them is defensible. None of them is recoverable from the number itself.
1. Which quantity is being named
CLIA lists the performance characteristics a laboratory must establish for a modified or in-house test, and analytical sensitivity is one of them.1 In CLIA usage that phrase does not mean the proportion of variants a test finds. It means the lowest amount of analyte the method can detect — the limit of detection.2 For an NGS assay that is a variant allele fraction and an input mass — 5% VAF at 20 ng of FFPE-derived DNA — not a percentage of anything found.
A figure like 99.5% is a different object: a detection rate, arrived at by counting. Both are called sensitivity in ordinary lab conversation, and a single validation summary can carry both without labelling either in a way that distinguishes them.
There is a third possibility, and in practice it is the most common one. The rate was almost certainly produced by comparing the new assay against something — a previous platform, an orthogonal method, a reference cell line, another laboratory’s result. FDA’s statistical guidance is unusually direct about what follows. When a test is evaluated against a comparator that is not a reference standard, sensitivity and specificity are not the appropriate terms; the identical arithmetic is reported as positive percent agreement and negative percent agreement.3 The guidance then adds the sentence worth pinning above the bench: agreement is not a measure of correctness. Two methods can agree closely and both be wrong.
This is not a preference about words. AMP, with liaison representation from CAP, frames the deliverable in exactly this vocabulary — positive percent agreement and positive predictive value, determined for each variant type.4 For an accredited laboratory that is the language the inspected document is expected to be written in, and a summary that says sensitivity where the study established agreement has quietly promoted its comparator to the status of truth. Sometimes it is defensible — a well-characterised reference material with an orthogonally confirmed answer key is close enough to truth for the purpose. The file should say so on the record rather than leave it implied by a word.
2. Over what denominator
Sensitivity is a fraction, and a fraction needs something underneath it. Three choices sit there, and each of them moves the number without anything about the assay changing.
Variant class. The AMP/CAP recommendation is to establish agreement and predictive value per variant type, and FDA asks for the same, calculated at both the variant and the specimen level.45 A blended figure is an average weighted by whatever mixture happened to be in the validation set — and in any such set, single-nucleotide variants dominate by count. A single number is therefore approximately the SNV number, with indel and copy-number performance folded invisibly into it. Nothing has been hidden. It has been averaged, which at the point of reading amounts to the same thing.
Territory. A detection rate is computed over a defined interval, and changing the interval changes the answer while the pipeline stands still. The cleanest demonstration in the literature is the Genome in a Bottle benchmark. Version 4.2.1 expanded the HG002 benchmark from 85% of GRCh38 to 92% of the autosomal assembly, mostly by adding segmental duplications and difficult-to-map regions — and scored against a short-read call set, it identified eight times more false negatives than the previous version.6 Same caller, same sample, same reads. The number moved because the denominator did. A sensitivity figure that does not travel with its interval file — and, where a reference sample was used, the version of the benchmark it was scored against — is a number without a scope.
No-calls. Positions that fail quality thresholds are neither successes nor failures; they are absences. FDA asks for the rate of no-calls and invalid calls to be estimated separately, with its own confidence interval, and treats an accuracy calculation performed after their removal as accuracy among valid results.5 Excluding them is legitimate. Not saying so converts a conditional claim — 99.5% of variants in positions the assay could read — into an unconditional one, and the gap between those two sentences is the entire uncovered fraction of the panel.
3. With what interval
The third choice is the one most often omitted altogether. A validation figure is an estimate from a finite set of observations, and the size of that set determines how much claim it can support.
The AMP/CAP guideline works this through explicitly: with 59 representative samples and no observed errors, a laboratory can state with 95% confidence that 95% or more of its samples will perform at or above the observed level.4 That is the shape of the statement the arithmetic actually supports — a bound, with a confidence attached, not a point.
FDA writes its own acceptance criterion the same way, as a pair rather than a number: a point estimate of no less than 99.9% with the lower bound of the 95% confidence interval no less than 99%, for all variant types.5 The regulator declines to accept a single figure, because a single figure does not say how much evidence is standing behind it.
On small validation sets the arithmetic is unforgiving. Recovering 199 of 200 known variants is 99.5%, with a 95% lower bound just above 97%. Recovering 995 of 1,000 is also 99.5%, with a lower bound of 98.8%. Two laboratories publishing the identical sentence are making claims that differ by nearly two points at the bound — and the bound, not the point estimate, is what a payer, a partner or an inspector is entitled to work from.
| Observed | Point estimate | 95% lower bound | What the file can defend |
|---|---|---|---|
| 59 / 59 | 100% | 93.9% | The AMP/CAP minimum: enough to state 95/95, not enough to claim two nines |
| 199 / 200 | 99.5% | 97.2% | Fails an FDA-style criterion on the bound while passing it on the headline |
| 995 / 1,000 | 99.5% | 98.8% | Same sentence, materially stronger claim — the difference is the n |
Wilson score intervals; the pattern matters, not the third decimal. The second row is the ordinary case — a laboratory can clear a 99.5% headline and still be unable to support a 99% floor — and if the acceptance criterion was written after the data arrived, that distinction never surfaces.
The number is not the problem
None of this argues that a headline figure is misleading, or that a laboratory should stratify its website into fourteen rows. A single number is a reasonable summary, and summaries are supposed to lose information. The failure mode is narrower: the summary is fine for exactly as long as the file underneath it can reconstruct what it meant, on request, without a reconstruction project.
Two things make that harder than it sounds. The first is that a modern NGS test has a bioinformatics pipeline inside it with performance characteristics of its own — so the file has to record which figures were established end-to-end with the wet-lab process and which against in-silico or pipeline-only data.7 The second is that the reconstruction already has an audience: CLIA requires a laboratory to make its established performance specifications available to clients on request.8 The stratified claim is not an internal artefact kept for inspection week. It is a document somebody outside the laboratory has a standing right to ask for — and the first to ask is usually the client deciding whether to send you samples.
What goes in the file
- A metric definition for every figure stated publicly: which quantity it is (limit of detection, detection rate, percent agreement), the named comparator, and whether that comparator is being treated as a reference standard and why.
- A stratified performance table, one row per variant type and specimen type, each carrying its numerator and denominator rather than a percentage alone.
- The interval files themselves — the target regions the denominator was computed over, versioned, plus the reference material and benchmark version where used.
- A no-call accounting over the same denominator, stating explicitly whether no-calls were excluded from the accuracy calculation or counted as misses.
- Confidence bounds and sample size per row, plus the acceptance criteria — dated, and demonstrably written before the data existed.
- A client-facing extract of the above, ready to send under §493.1291(e) without assembling it from scratch each time.
One test for whether the file is in that state: take the sensitivity figure currently on the laboratory’s website and try to recover, in ten minutes, which comparator produced it, which variant classes and intervals it covers, and how many observations sit behind it. If that takes an afternoon, the number is not yet a claim the laboratory owns. It is one the laboratory is hosting.
References
- 42 CFR §493.1253 — Standard: Establishment and verification of performance specifications. Electronic Code of Federal Regulations. ecfr.gov
- Association of Public Health Laboratories. CLIA Inspection Guidance Document (companion to the CLIA Interpretive Guidelines), 2013 — analytical sensitivity as the lowest concentration the test can distinguish from a blank, usually termed the limit of detection. aphl.org
- US Food and Drug Administration. Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests. Guidance for Industry and FDA Staff, 2007. fda.gov
- Jennings LJ, Arcila ME, Corless C, et al. Guidelines for Validation of Next-Generation Sequencing–Based Oncology Panels: A Joint Consensus Recommendation of the Association for Molecular Pathology and College of American Pathologists. J Mol Diagn. 2017;19(3):341–365. doi:10.1016/j.jmoldx.2017.01.011
- US Food and Drug Administration. Considerations for Design, Development, and Analytical Validation of Next Generation Sequencing (NGS)-Based In Vitro Diagnostics (IVDs) Intended to Aid in the Diagnosis of Suspected Germline Diseases. Guidance for Stakeholders and FDA Staff, April 2018. fda.gov
- Wagner J, Olson ND, Harris L, et al. Benchmarking challenging small variants with linked and long reads. Cell Genomics. 2022;2(5):100128. doi:10.1016/j.xgen.2022.100128
- Roy S, Coldren C, Karunamurthy A, et al. Standards and Guidelines for Validating Next-Generation Sequencing Bioinformatics Pipelines: A Joint Recommendation of the Association for Molecular Pathology and the College of American Pathologists. J Mol Diagn. 2018;20(1):4–27. doi:10.1016/j.jmoldx.2017.11.003
- 42 CFR §493.1291(e) — Standard: Test report; performance specifications made available to clients upon request. Electronic Code of Federal Regulations. ecfr.gov

