Reading a Validation Report: What Analytical Validation Establishes, and Why Clinical Validity Is a Separate Document

Reading a Validation Report: What Analytical Validation Establishes, and Why Clinical Validity Is a Separate Document
Zetobit Reading the Report
READING THE REPORT Reading a Validation Report What it establishes, and why clinical validity is a second document Kanna Nandakumar, PhD Zetobit
Reading the Report

Reading a Validation Report: What Analytical Validation Establishes, and Why Clinical Validity Is a Separate Document

The numbers describe how closely the assay agrees with a comparator, on specimens the laboratory assembled on purpose. Whether the thing it measures says anything about your patient is a different question, settled by a different body of evidence.

A validation report is the document a laboratory produces to show that its test works, and it is usually convincing, because it is full of high numbers with tight confidence intervals. The numbers are almost always correct. What is easy to miss is what they are numbers about. A validation report describes an instrument, measured against another instrument, on a set of specimens chosen to make that measurement possible. It is not a statement about your patient, and it does not claim to be one.

An earlier piece in this series treated validation as one of several things to ask about a pipeline nobody had named. Here the report is in front of you, and the question is narrower: what do its contents actually establish, and what would you need to read next?

01The document was written to a standard that stops short on purpose

The performance characteristics a laboratory must establish before reporting patient results are set out in the CLIA regulations: accuracy, precision, analytical sensitivity, analytical specificity including interfering substances, reportable range, reference intervals, and any other characteristic required for the test to perform.1 Every item on that list is a property of the measurement. None of them is a property of the disease.

The regulator says this outright. In joint guidance with the FDA, CMS states that its CLIA program does not address the clinical validity of any test — the accuracy with which the test identifies, measures, or predicts the presence or absence of a clinical condition or predisposition in a patient.1 The agency that certifies the laboratory is telling you, in advance, which question its certification does not answer.

That is a division of labor rather than a loophole. Analytical validity is a property of an assay in one laboratory’s hands and can reasonably be established there, once. Clinical validity is a property of the relationship between a measurement and a condition across a population, and no single laboratory can establish it on ninety-six specimens no matter how carefully it works.

02Sensitivity is measured against a comparator, not against a disease

An analytical sensitivity figure answers one question: of the variants the reference method said were present, what fraction did this assay recover? The comparator is the boundary of the claim. Anything the reference method could not see is outside the measurement entirely, and any class of variant absent from the validation set has no measured sensitivity — not a low one, an absent one. The AMP/CAP consensus recommendations make the scope explicit, requiring that validation be appropriate to the intended clinical use, the specimen types, and the variant types of the test.2 So the informative part of the document is the cohort composition table, not the headline percentage: how many specimens, of which types, containing how many of each class of variant, confirmed by what.

It is worth being pedantic about one word. Analytical sensitivity is the fraction of present variants detected. Clinical sensitivity is the fraction of affected patients identified — which depends on how much of the disease is caused by mechanisms the assay can see at all. An assay with flawless analytical sensitivity across its content has whatever clinical sensitivity its content happens to confer, and that number does not appear in the validation report.

VALIDATION SUMMARY — EXCERPT Validation cohort 96 specimens Reference method orthogonal assay Analytical sensitivity 99.4% Analytical specificity >99.99% Positive predictive value98.7% Reproducibility 100% concordant This is the denominator. What it missed is invisible. Inherits the cohort’s prevalence, not your population’s. THIS DOCUMENT ANSWERS Does the assay recover what the comparator says is there? Does it give the same answer twice? Over what range, in which specimen types, for which variant classes? Established by the laboratory. Required by CLIA. A DIFFERENT DOCUMENT ANSWERS Does what it measures relate to the condition being asked about? How strong is that gene–disease link? What fraction of test positives are affected in this population? Built from literature, curation, and cohort studies. THE DECOUPLING A perfectly validated assay can report a technically correct result in a gene whose disease association does not hold.

Each figure in a validation summary is a statement about a different object. Sensitivity and specificity describe the assay against its comparator and carry across populations; the predictive value describes the cohort and does not. Nothing on the page speaks to whether the measured thing bears on the condition.

03The cohort was assembled, so the predictive values do not transfer

This is the most consequential line in most validation reports and the least discussed. Sensitivity and specificity are properties of the test and do not depend on how common the condition is. Predictive values are not: they depend on prevalence. And a validation cohort is deliberately enriched for positives, because you cannot measure sensitivity without them — which means any predictive value computed inside it is a fact about the cohort rather than about practice.

A published validation of cell-free DNA screening for a recurrent microdeletion makes the gap unusually visible. Positive predictive value observed in the study population was 100%; in the same paper, the expected PPV in the general pregnant population was estimated at 12.2% at 99.5% specificity.3 A later prospective study with genetic confirmation in more than 18,000 pregnancies, where the condition occurred about once in 1,500, reported sensitivity of 75.0%, specificity of 99.84%, and a positive predictive value of 23.7%.4 Nothing malfunctioned. A specificity of 99.84% is excellent. It is simply that when a condition is rare, even a very small false-positive rate generates more false positives than true ones — and the validation cohort, by construction, was not rare.

The practical rule is short. If a validation report quotes a predictive value, find out what prevalence it assumed and whether that matches the population you intend to test. If the document does not say, the number is a description of the sample set.

04Clinical validity is assessed elsewhere, by a different method

Clinical validity has its own evidence framework, its own curation process, and its own published classifications — and the two can come apart completely. When a ClinGen expert panel applied that framework to 142 genes drawn from the panels of 17 clinical laboratories, 45 of the 164 gene–disease pairs assessed came out Limited, Disputed, or Refuted.5 Every one of those genes was already being offered clinically, and every one of them could have been sequenced with impeccable analytical accuracy. The base calls would have been right. The inference would not have been.

The professional response is explicit about which document governs. The ACMG technical standard on diagnostic gene panels advises that laboratories use gene–disease validity classifications when designing panels and report clinically only on genes supported by at least moderate evidence, with variants in genes carrying only limited evidence not classified above uncertain significance.6 That is not a comment on assay performance. It is an instruction to consult a second body of evidence before treating a correct measurement as a finding.

What to do with the document you have

A validation report is necessary and it is not sufficient, and asking for the second document is not a criticism of the first. For a well-established indication, the clinical validity evidence already exists in the literature and the guidelines, and a good laboratory will point you to it without hesitation — often it is summarized in a page of the same packet. For a novel indication, a newly added gene, or a population unlike the one studied, someone has to have done the work, and it is worth establishing who did it and in whom.

The confusion has a predictable direction. Analytical figures are the ones that are straightforward to generate, easy to state, and flattering, so they are the ones that travel — into brochures, investor decks, and diligence summaries, where a sensitivity of 99.4% is read by everyone in the room as a claim about patients. Nobody has lied. A number computed against an orthogonal assay on ninety-six banked specimens has simply been carried, unlabeled, into a conversation about clinical performance, and once it is there it is difficult to put back.

The failure mode here is rarely a bad validation report. It is a good one, read as though it answered a question it was never designed to answer — and read that way, its high numbers and tight intervals make it more persuasive, not less.

Three questions to ask

  1. Which specimen types and classes of variant were in the validation cohort, and what reference method were they measured against?
  2. If a predictive value is quoted, what prevalence does it assume — and does that match the population I am testing?
  3. Where is the evidence that what this assay measures relates to the condition I am asking about, and is it in this document or another one?

References

  1. Centers for Medicare & Medicaid Services and US Food and Drug Administration. FDA and CMS: Laboratory Developed Tests — Frequently Asked Questions; and 42 CFR §493.1253, Standard: Establishment and verification of performance specifications. cms.gov
  2. Roy S, Coldren C, Karunamurthy A, et al. Standards and Guidelines for Validating Next-Generation Sequencing Bioinformatics Pipelines: A Joint Recommendation of the Association for Molecular Pathology and the College of American Pathologists. J Mol Diagn. 2018;20(1):4–27. doi:10.1016/j.jmoldx.2017.11.003
  3. Bevilacqua E, Jani JC, Chaoui R, et al. Performance of a targeted cell-free DNA prenatal test for 22q11.2 deletion in a large clinical cohort. Ultrasound Obstet Gynecol. 2021;58(4):597–602. doi:10.1002/uog.23699
  4. Dar P, Jacobsson B, Clifton R, et al. Cell-free DNA screening for prenatal detection of 22q11.2 deletion syndrome. Am J Obstet Gynecol. 2022;227(1):79.e1–79.e11. ajog.org
  5. DiStefano MT, Hemphill SE, Oza AM, et al. ClinGen expert clinical validity curation of 164 hearing loss gene–disease pairs. Genet Med. 2019;21(10):2239–2247. doi:10.1038/s41436-019-0487-0
  6. Bean LJH, Funke B, Carlston CM, et al. Diagnostic gene sequencing panels: from design to report — a technical standard of the American College of Medical Genetics and Genomics (ACMG). Genet Med. 2020;22(3):453–461. doi:10.1038/s41436-019-0666-z
Reading the Report is a Zetobit series for people who receive genomic and multi-omic work rather than produce it. Each piece takes the document as given and asks what it establishes, what it leaves open, and what to ask next.
Previous
Previous

Diligence on a Genomics Claim: What to Ask About the Cohort, the Split, and the Endpoint

Next
Next

Reading a Bioinformatics Scope of Work: What’s In, What’s Out, and What Gets Billed Later