Inter-Laboratory Discordance
Inter-Laboratory Discordance
Why the Same Variant Gets Different Answers on the Same Day
A pathogenicity classification looks like a property of the variant. It is written as one — a single word in a report, attached to a genomic coordinate. It is closer to a property of the laboratory that issued it, and the gap between those two readings is measurable.
An earlier piece in this series covered classification drift along the time axis: the same laboratory, the same variant, a different answer years later as evidence accumulated. That kind of change is defensible — it is what a working evidence framework is supposed to do. This piece covers the other axis. Same variant, same day, same published guidelines, different laboratories. Nothing has changed except who is answering.
Structurally this is the MSI and TMB piece applied one layer up. That piece argued that a number is a property of the assay that produced it, not of the tumour. The same argument holds for a category, with an important difference: two assays measuring MSI genuinely measure different things — different loci, different depths, different algorithms. Here every laboratory is looking at the same variant, and most of them at the same public evidence. The disagreement has to come from somewhere else.
The size of the problem, and how it moved
The foundational measurement came a year after the ACMG/AMP guidelines were published. Nine laboratories in the CSER consortium classified 99 variants using both the new guidelines and their own in-house methods. Within a laboratory, the two methods agreed 79% of the time. Across laboratories, concordance was 34% for either system — and after consensus discussion and detailed review of the criteria, it rose to 71%1.
Two findings in that study are easy to skip past and worth stating plainly. First, adopting a shared framework did not, by itself, improve agreement: there was no statistically significant difference in inter-laboratory concordance between classifications made with the ACMG/AMP criteria and those made with each laboratory's own rules1. A common vocabulary is not a common answer. Second, the framework has a direction: in 48 of 72 discordant calls (67%), the ACMG/AMP result sat closer to VUS than the laboratory's own method did1.
A follow-up four years later gives the trajectory. Eight laboratories each submitted 20 classified variants in the ACMG secondary-findings genes; each of the resulting 158 variants was independently reclassified by two further laboratories, blinded to the original classification and evidence codes. Complete five-category concordance across three laboratories was 54%, with 11% showing a discordance that could affect clinical recommendations — and no variant ranged all the way from P/LP to LB/B, against 5% in the original study2. After review, concordance reached 84%2.
Real improvement, and the authors are careful to note the earlier 34% understated general concordance because that variant set was enriched for difficult cases2. But 54% complete agreement on a randomly distributed set, from laboratories using the same guidelines on the same genes, is still the number to hold in mind when reading a single report.
Where the disagreement actually comes from
The most useful study here is not a concordance measurement but a diagnosis. Four clinical laboratories compared their ClinVar submissions: of 6,169 variants classified by at least two of them, 88.3% were initially concordant. They then took a subset of the differences, documented the basis for each, shared internal data, and independently reassessed3.
The breakdown is the finding. Over half of the differences resolved for reasons that are not interpretive disagreement at all: 36% because reassessing an old classification with the laboratory's own current criteria changed it, and 17% because the reinterpretation had already been done but not yet submitted to ClinVar. Differences in internal data accounted for 33% — segregation (10%), co-occurrence (9%), internal proband frequency (8%), detailed phenotype (6%). Differences in the use or weighting of public data accounted for 14%, including benign thresholds (9%) and different data sources (5%)3.
That decomposition reframes the problem. The headline — laboratories disagree about variants — is mostly not what is happening.
The largest band is the temporal axis observed from outside. When a laboratory's ClinVar entry reflects criteria it no longer uses, a comparison between two laboratories is partly a comparison between two different points in time. The drift piece and this one turn out to describe the same mechanism from different vantage points: what looks like inter-laboratory disagreement is substantially intra-laboratory staleness, made visible by putting two snapshots side by side.
The second band is not disagreement in any sense — the laboratories already agreed and the database had not caught up.
The third is the one that genuinely cannot be resolved by better rules, and it is not a rules problem at all. It is an information asymmetry. One laboratory has seen the variant segregate in a family; another has not. One has three internal probands with a matching phenotype; another has none. Both applied the guidelines correctly to the evidence in front of them, and reached different answers because the evidence in front of them differed.
Only the last 14% is what most people picture: the same evidence, read differently. It is the smallest band, and the only one that clearer guidelines by themselves could address.
Likely pathogenic is the unstable category
The discordance is not spread evenly across the five categories, and the pattern is consistent enough to be actionable.
In the 2020 study, 21% of variants submitted as pathogenic were discordant with VUS — but 63% of variants submitted as likely pathogenic were2. Concordance was highest for variants submitted as VUS, which the authors attribute to a common lack of prior observations, leaving little evidence for anyone to weigh differently2. Agreement is easiest where there is nothing to disagree about.
The instability of LP has a structural cause. The ACMG/AMP framework defines five categories5, and when it is modelled as a Bayesian classifier, likely pathogenic corresponds to a posterior probability of pathogenicity of roughly 90–99% and VUS to 10–90%6. Evidence arrives in discrete increments — supporting, moderate, strong — and combinations of increments do not land neatly on either side of a 90% line. A single evidence code applied by one laboratory and not another moves a variant across the boundary. For P, the evidence is usually overwhelming enough that one code does not decide it.
Practically: an LP call is the classification most likely to be something else at the laboratory down the road, and it is also a classification that in most protocols triggers the same clinical action as P.
| Axis | What differs | Is the change legitimate? | Remedy |
|---|---|---|---|
| Temporal (drift) | Evidence available at two dates | Yes — this is the framework working | Scheduled reanalysis |
| Laboratory, stale record | One party's criteria are older | No — an artifact of update cadence | Re-date and resubmit |
| Laboratory, private evidence | Internal cases one party holds | Both correct on what they can see | Data sharing; nothing else works |
| Laboratory, weighting | Codes applied to identical evidence | Genuine interpretive difference | Gene-specific criteria; point systems |
| Assay (the MSI/TMB case) | The underlying measurement itself | Different quantities, not different opinions | Report the assay with the number |
The rows demand different responses, which is why collapsing them into “labs disagree” is unhelpful. Only the fourth row is an interpretation problem in the ordinary sense; the third is a structural feature of a field where the most informative evidence sits in private case records.
What has actually reduced it
Three interventions have measurable effects, and it is worth being clear about which does what.
Comparing and discussing. This is the most reliable lever in every study cited here: 34% to 71% after consensus discussion1, 54% to 84% after review2, and 87.2% of reassessed discordant variants resolved in the four-laboratory ClinVar project, which lifted overall concordance from 88.3% to 91.7%3. The authors' framing is worth keeping: sharing interpretations allows differences to be identified and creates the motivation to resolve them3.
Specifying the criteria for a gene or disease. The generic guidelines were written to apply everywhere, which means the thresholds inside them are unspecified for anywhere in particular. ClinGen's expert panels exist to fix that, and the 2020 study points directly at which codes need it — PM1 for mutational hotspots and PS4 for case-control data, particularly in relation to applying PM2 for variants absent from controls2.
Converting the rules into arithmetic. A points-based implementation removes the ambiguity of combining codes by hand. The clearest demonstration comes from the copy-number side: nine laboratories classifying 234 distributed CNVs went from 18% complete concordance under their previous methods to 76% using the ACMG/ClinGen scoring metric, reaching 85% after review, with 11% of classifications showing discordance that could affect medical management4. A four-fold improvement from making the combination step mechanical.
The parallel for sequence variants is the Bayesian points framework now embedded in ACMG/AMP practice6. Note what it does and does not fix: it standardises how evidence is combined, not which evidence each laboratory has. The 33% band is untouched by it.
What this means for a report in front of you
A classification is a statement of the form: given this evidence, weighted this way, on this date, we conclude X. Reports render it as the variant is X. The compression is the problem, and it is not the laboratory's fault — a clinician needs an actionable category, not a probability distribution over five of them.
But the compression means the reader cannot see which of the four situations they are in. A P/LP versus VUS discordance affecting clinical recommendations occurred for 11% of variants in the 2020 study2 and 11% of CNVs in the copy-number study4 — two different variant types, two different laboratory networks, the same order of magnitude.
What to do about it
- Read the evidence codes, not just the category. The codes show which criteria carried the call. A classification resting on one moderate code is a different object from one resting on three strong codes, and only the first is likely to differ elsewhere.
- Check ClinVar for other submitters and read the dates. Given that stale records were the largest single cause of apparent disagreement3, a conflict where one submission is years older may not be a conflict.
- Treat LP with more caution than P. It is the category most likely to be VUS at another laboratory2, while usually triggering the same clinical action.
- Prefer a ClinGen expert-panel classification where one exists. It reflects gene-specific criteria and pooled evidence rather than one laboratory's view.
- Submit your own classifications, with evidence. The mechanism that resolves discordance is comparison, and comparison requires both parties to have published. A laboratory that consumes ClinVar without contributing benefits from the process while withholding the input it depends on.
- When you get a second opinion, ask what internal evidence it rests on. If the two laboratories differ because one has segregation data, that is not a tie to be broken by preference — the additional evidence is simply better.
- Use a points-based implementation so that combining evidence is arithmetic rather than judgement4, and record the points, not only the outcome.
- Record the classifying laboratory, its criteria version and the date alongside the call — the same discipline the drift piece argued for, extended to identity as well as time.
- Do not average across laboratories. A variant with two LP calls and one VUS is not “mostly LP”; one submitter may hold decisive evidence the others lack. Read the reasons, not the vote.
The shape of the error
Most entries in this series describe a measurement carrying an inference it cannot support. This one is different in an instructive way: there is no measurement error anywhere. The sequencing is right, the variant call is right, the guidelines were applied competently by every party. And three laboratories still produce three reports.
The reason is that a classification was never a measurement. It is a conclusion drawn from an evidence set, and evidence sets differ between institutions in ways that have nothing to do with competence — one has more cases, another curated more recently, a third weights population frequency slightly differently. The ACMG framework standardises the reasoning. It cannot standardise the inputs.
Which suggests the useful question when two reports disagree is not which laboratory is right. It is what does each one know that the other does not — because in a third of documented cases, that question has a concrete answer, and it is not a matter of opinion at all.
References
- Amendola LM, Jarvik GP, Leo MC, et al. Performance of ACMG-AMP variant-interpretation guidelines among nine laboratories in the Clinical Sequencing Exploratory Research Consortium. American Journal of Human Genetics 2016;98(6):1067–1076. cell.com/ajhg — S0002-9297(16)30059-3
- Variant classification concordance using the ACMG-AMP variant interpretation guidelines across nine genomic implementation research studies. American Journal of Human Genetics 2020;107(5):932–941. Author list not captured — verify before publication. cell.com/ajhg — S0002-9297(20)30356-6
- Harrison SM, Dolinsky JS, Knight Johnson AE, et al. Clinical laboratories collaborate to resolve differences in variant interpretations submitted to ClinVar. Genetics in Medicine 2017;19(10):1096–1104. nature.com/articles/gim201714
- Adaptation of ACMG-ClinGen technical standards for copy number variant interpretation concordance. Frontiers in Genetics 2022;13:829728. Author list not captured — verify before publication. frontiersin.org — fgene.2022.829728
- Richards S, Aziz N, Bale S, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genetics in Medicine 2015;17(5):405–424. nature.com/articles/gim201530
- Tavtigian SV, Greenblatt MS, Harrison SM, et al., on behalf of the ClinGen Sequence Variant Interpretation Working Group. Modeling the ACMG/AMP variant classification guidelines as a Bayesian classification framework. Genetics in Medicine 2018;20(9):1054–1060. nature.com/articles/gim2017210

