How Many Positives Do You Need: Why 25, 59 and 600 Are All Correct Answers

How Many Positives Do You Need: Why 25, 59 and 600 Are All Correct Answers — Zetobit
Zetobit  /  Clinical Bioinformatics
POSITIVES THE VALIDATION FILE How Many Positives? Three documents. Three minimums. No contradiction. 25 59 600 Kanna Nandakumar, PhD  ·  Zetobit

The Validation File

How Many Positives Do You Need: Why 25, 59 and 600 Are All Correct Answers

Three authoritative documents give three different minimums for the same assay. They do not disagree. Each counts a different object toward a different obligation.

It is the most common question asked before a validation begins, and the one with the least satisfying answer. How many positive samples do we need? Somebody has heard 59. Somebody else has a state guideline open to a page that says 25. A third remembers a laboratory that kept confirming variants orthogonally for two years after going live. All three are describing something real, and why they can all be right at once is the subject of this piece.

The previous piece in this series took apart what a sensitivity figure is a claim about. This one takes the next step: how many observations you need before you are entitled to make the claim at all. The answer turns on an object that most validation plans never name — not the confidence interval around an estimate, but the tolerance interval around a distribution.

Start with the silence that creates the problem. CLIA lists what a laboratory must establish for a test it has modified or developed — accuracy, precision, analytical sensitivity and specificity, reportable range, reference intervals — and names no sample count for any of them.1 That is not an oversight; a rule governing every method in every discipline cannot sensibly specify one. But a silence in a regulation gets filled, and in clinical NGS it has been filled three times over, by three bodies answering three different questions.

The 59 that everybody quotes

AMP's oncology panel validation guideline, developed with liaison representation from CAP, recommends a minimum of 59 samples.2 The figure is quoted constantly and misread almost as often, because most people assume it is a sample-size calculation for sensitivity. It is not.

59 is the smallest n at which a one-sided nonparametric tolerance interval reaches 95% reliability at 95% confidence with zero failures. The guideline is careful about the distinction: a confidence interval speaks to the average performance of a population, a tolerance interval to the performance of an individual sample.2 Its worked example makes that concrete — a director who runs 59 representative samples and sees a worst-case false-positive rate of 1.9% can state with 95% confidence that at least 95% of samples will fall at or below 1.9%. That is a claim about the worst case you should expect next month, not about your mean.

Two riders travel with the number and get dropped more often than it does. The 59 applies per variant type — SNVs, small indels, large indels, copy number alterations and structural variants each carry their own count. And where a laboratory cannot source 59 for a variant type, the guideline's remedy is not a smaller number; it is to supplement the NGS test with another validated method.2 The shortfall does not become a lower bar. It becomes a second assay.

Where 59 comes from, and what moves it

The arithmetic is one line. With zero failures, you need the smallest n for which reliability raised to the power of n falls at or below one minus your confidence level. At 95% confidence, that gives 29 samples for 90% reliability, 59 for 95%, 299 for 99%, and 2,995 for 99.9%.

Zero-failure sample counts at 95% confidence
Reliability you want to claimn, no failuresn, one failureWhat the sentence sounds like
90%2946Rarely defensible for a clinical variant class
95%5993The quoted minimum
99%299473A 99% floor, per variant type
99.9%2,9954,742Not reachable from patient accrual

Two consequences follow, and the second costs money. A paired criterion of the kind regulators write for high-stakes claims — a point estimate at or above 99.9% together with a lower bound at or above 99% — cannot be met by 59 of anything; a 99% floor needs roughly 299 clean observations of that variant type. And every count in the table assumes perfection. One missed variant among 59 does not cost you a fraction of a percentage point: to make the same 95/95 statement having seen a single failure, you need 93 samples. The price of one false negative is thirty-four more positives.

The 25 that New York asks to see

New York State's germline NGS guidelines ask for something smaller and different in kind: a minimum of 25 unique patient samples for initial validation, spanning every specimen type accepted, with a representative distribution of variant types across all targeted regions including GC-rich and homopolymer sequence, each confirmed by an independent reference method that may not use the same technology unless it runs in a different laboratory.3 Around that sit smaller counts for narrower questions — two well-characterized reference samples to establish a laboratory-specific error rate expected below 2%; three to five specimens per variant type for the input limit of detection; three positives per variant type across three runs on different days for reproducibility. The somatic guidelines follow the same logic on their own axis, asking for three samples in each tumor mutational burden class across three runs.4

25 is not a competing answer to 59. It is a gate — what a regulator requires to see before the assay goes on the permit. Critically, it is also not the point at which confirmation stops.

The 600 that almost nobody counts

The least-read part of the same document is where the largest number lives. After initial validation, New York requires continued orthogonal confirmation of reportable variants until stratified thresholds are reached.3 SNPs are divided into three sequence contexts — standard, GC-rich at 60% or above, and homopolymer runs longer than three bases — requiring 50, 75 and 75 confirmed calls respectively. Indels carry the same three-way split and the same counts. CNVs divide by size into single-gene events, whole-arm events and the megabase-scale range between them, requiring 100, 50 and 50. Six hundred orthogonally confirmed calls, in total, before an assay is free of routine confirmation.

And the count is not really a count. Each category caps how many confirmations may come from any one place: no more than five in a particular gene for SNPs and indels, no more than ten at a particular locus for single-gene CNVs — with the cap doubling as a retirement rule, so that once a gene has been confirmed five times, that gene no longer needs confirmation before reporting. Fifty standard-context SNPs at five per gene cannot come from fewer than ten distinct genes. The number is a proxy for something the document cannot easily specify directly, which is coverage of contexts. The cap is the actual requirement. The total is what the cap implies.

Set the three side by side and the apparent conflict dissolves. 59 answers what may I claim about a sample I have not run yet. 25 answers what must a regulator see before go-live. 600 answers when may I stop confirming. A single assay can owe all three at the same time, and satisfying one says nothing about the others.

One variant type × one sequence context Claim about a future sample Tolerance interval, 95 / 95 Claim of readiness to report Initial validation set, all specimen types Claim that confirmation can stop Stratified counts, capped per gene 59 zero failures, per variant type 25 unique patients, orthogonally confirmed 50–100 per bucket · 600 across all DELIVERABLE Tolerance-interval statement DELIVERABLE Validation summary table DELIVERABLE Confirmation-retirement log Not enough physical positives exist Reference material characterized, traceable Verified in silico set record the file level used A second validated method orthogonal, standing A standing line item in the SOP — never a smaller number
The same assay owes three different counts because it is making three different claims. Where positives cannot be sourced, every framework routes the shortfall into a documented substitute rather than a reduced threshold.

The pipeline gets a different denominator

For a bioinformatics pipeline, the accrual constraint is softer than it looks. The AMP/CAP/AMIA pipeline standards defer the minimum-count question to the oncology guideline, then add that where the wet-laboratory minimum cannot exercise the pipeline across the variant types and allele fractions it must handle, reference materials and verified in silico samples may be analyzed as well.5 A later joint report from AMP, the Association for Pathology Informatics and CAP catalogs the mechanisms: reads simulated from reference sequence, two samples mixed to hit a target allele fraction, a downsampled FASTQ, variants injected into real laboratory files, and reanalysis of existing data to test a pipeline change.6

This is a real escape from sample scarcity, and it carries a trap worth writing into the plan rather than discovering during an inspection. Variants introduced into an aligned BAM never pass through the aligner. Unless the mutagenized file is converted back to FASTQ, the exercise tests a shorter pipeline than the one producing patient results — and alignment is precisely where medium indels and repetitive contexts are lost. Purely simulated reads have the mirror problem: they carry any variant you specify but cannot reproduce every bias of the real assay. An in silico positive counts toward whatever it actually exercised, which makes the honest record a statement about file level, not a tally of samples.

What the three numbers have in common

None of these frameworks lets scarcity reduce the obligation. AMP and CAP convert a shortfall into a second validated method. New York converts it into a confirmation program that runs until the count is reached, gene by gene, locus by locus. The pipeline standards convert it into a documented in silico set with its limitations stated. The federal regulation, naming no count at all, leaves the laboratory director to write one down and defend it.1 In every case the number you cannot reach does not vanish when a smaller one is written in its place; it reappears as a standing line item in the SOP. An inspector who finds neither the count nor its substitute is looking at a performance claim with nothing underneath it.7

Which makes the original question the wrong shape — not answerable as asked, in the way that how long is the rope is not answerable without knowing what it has to reach. How many of what, in which sequence context, to support which sentence, at what confidence: a laboratory that can put that in a table before collecting a single sample has already done the difficult part. The numbers fall out of the table. They were never the decision.

What goes in the file

  • A validation plan dated before data collection that states, for each variant type and sequence context, the target reliability, the confidence level, and the minimum count those two imply — with the arithmetic shown, not just the result.
  • A per-variant-type inventory: variant class, context bucket, number of positives, source of each (patient specimen, reference material, in silico), orthogonal method used, and outcome.
  • An explicit statement of which claim each count supports — a tolerance interval on future per-sample performance is a different sentence from a confidence interval on an estimate, and the file should say which one you are making.
  • For every variant type falling short of target: the named substitute — supplementary validated method, reference material, or ongoing confirmation policy — together with its retirement criterion in writing.
  • For in silico positives: the generation method, the file level at which variants were introduced, and a plain statement of which pipeline stages were therefore not exercised.
  • The reproducibility record as a grid: positives per variant type × independent runs × days × operators × instruments.
  • A revalidation trigger list mapping each class of change — software version, reagent lot, reference build — to the specific subset of the validation set that must be rerun.

References

  1. 42 CFR § 493.1253 — Standard: Establishment and verification of performance specifications. Electronic Code of Federal Regulations. ecfr.gov
  2. Jennings LJ, Arcila ME, Corless C, Kamel-Reid S, Lubin IM, Pfeifer J, Temple-Smolkin RL, Voelkerding KV, Nikiforova MN. Guidelines for Validation of Next-Generation Sequencing–Based Oncology Panels: A Joint Consensus Recommendation of the Association for Molecular Pathology and College of American Pathologists. J Mol Diagn. 2017;19(3):341–365. jmdjournal.org
  3. New York State Department of Health, Clinical Laboratory Evaluation Program. Next Generation Sequencing (NGS) Guidelines for Germline Genetic Variant Detection. Updated and revised, October 2024. wadsworth.org
  4. New York State Department of Health, Clinical Laboratory Evaluation Program. Next Generation Sequencing (NGS) Guidelines for Somatic Genetic Variant Detection. wadsworth.org
  5. Roy S, Coldren C, Karunamurthy A, Kip NS, Klee EW, Lincoln SE, Leon A, Pullambhatla M, Temple-Smolkin RL, Voelkerding KV, Wang C, Carter AB. Standards and Guidelines for Validating Next-Generation Sequencing Bioinformatics Pipelines: A Joint Recommendation of the Association for Molecular Pathology and the College of American Pathologists. J Mol Diagn. 2018;20(1):4–27. sciencedirect.com
  6. Duncavage EJ, Coleman JF, de Baca ME, Kadri S, Leon A, Routbort M, Roy S, Suarez CJ, Vanderbilt C, Zook JM. Recommendations for the Use of in Silico Approaches for Next-Generation Sequencing Bioinformatic Pipeline Validation: A Joint Report of the Association for Molecular Pathology, Association for Pathology Informatics, and College of American Pathologists. J Mol Diagn. Published online 13 October 2022. doi:10.1016/j.jmoldx.2022.09.007. jmdjournal.org
  7. Centers for Disease Control and Prevention, Division of Laboratory Systems. NGS Quality Initiative — additional materials and standards landscape. cdc.gov
Previous
Previous

What Counts as the Same Event: Validating Copy Number and Structural Variant Detection

Next
Next

Reportable Range Is Not the Gene List