MSI and TMB Across Assays: Why the Same Tumor Yields Different Numbers
MSI and TMB Across Assays: Why the Same Tumor Yields Different Numbers
Two laboratories receive tissue from the same block. One reports a tumor mutational burden of 11 mutations per megabase; the other reports 6. Neither made an error. TMB is not a physical property being measured with differing precision — it is a quantity constructed from a definition, and the definition includes which territory was sequenced, which variants were counted, how germline was removed, and what the denominator was. Change any of those and the number changes, legitimately. What reaches the report is a scalar and a threshold.
Most measurements in clinical genomics have a referent. A variant is present or it is not. A gene is deleted or it is not. Reproducibility means two laboratories converging on the same underlying fact, and when they disagree, one of them is wrong.
TMB does not work this way, and neither does an NGS-derived MSI score. Both are summary statistics computed over a chosen territory under a chosen set of counting rules. There is a true underlying quantity — the tumor's actual somatic mutation rate, the tumor's actual mismatch repair status — but the reported number is an estimator of it, and estimators differ in ways that are not error.
This matters because both numbers are compared against fixed thresholds that carry treatment decisions. Pembrolizumab's tissue-agnostic indication uses TMB ≥ 10 mut/Mb. MSI-high status gates checkpoint inhibitor eligibility across tumor types. A threshold implies a standard scale, and the scale is what varies.
The series covered a related problem recently: FFPE deamination artifacts inflating TMB and disturbing MSI calls. That is a specimen effect — the same assay applied to differently-preserved tissue. This piece is the other axis: the same tissue analyzed under different assay definitions. The two compound, but they are separate failures, and only one of them is fixed by better tissue handling.
The consortium result, and why it is reassuring and alarming at once
The Friends of Cancer Research TMB Harmonization Project was formed specifically because of this problem, with participation from assay developers, pharmaceutical companies, the FDA, and academic groups.1 Its second phase distributed 29 tumor samples and 10 human-derived cell lines to 16 laboratories, each of which ran its own bioinformatics pipeline and compared results against whole exome sequencing.2
The finding was that estimation of TMB varies across panels, with panel size, gene content, and bioinformatics pipeline all contributing to empirical variability.2 A calibration approach derived from TCGA data, tailored per assay, reduced the spread of panel values around the WES value for 26 of 29 clinical samples.2
Read that last sentence carefully, because it contains the whole argument. Calibration worked — which means the divergence was systematic and correctable, not random noise. Each assay was measuring something real and doing so consistently. They simply were not measuring the same thing, and the differences were stable enough to be modelled and removed.
That is the situation the field is in: the numbers are individually valid, mutually incomparable, and reconcilable only if you know the assay parameters — which the report does not carry.
Panel size: the arithmetic before any biology
Before considering counting rules, there is a purely statistical floor. TMB is a rate estimated from a sample of the genome, and the precision of a rate estimate depends on how much was sampled.
The relationship has been characterized precisely: the coefficient of variation of panel-based TMB decreases inversely with the square root of the panel size and the square root of the TMB level. In silico simulation across commercially available panels in the TCGA pan-cancer cohort confirmed this and found a CV of 35% at TMB = 10 mut/Mb even for the largest panels of 1.1–1.4 Mb.3
The consequence for a single patient is best stated the way the authors state it. For a patient whose TMB is estimated by a 2 Mb panel, a given CV corresponds to a confidence interval of roughly 10.1–21.4 mut/Mb. Run the same patient on a 1 Mb panel and the interval becomes about 8.4–24.7; on a 0.5 Mb panel, about 6.3–30.2.3 The point estimate can sit comfortably above the threshold while the interval spans it.
Independent work reaches the same place from the Bayesian direction: with only 0.3 Mb sequenced, an observed TMB of 12 mut/Mb implies roughly an 81% chance that the true TMB exceeds 10 — meaning close to a one-in-five chance the patient is above the line only by sampling.4
Consortium work converged on a lower bound: panel sizes above 667 Kb are necessary to maintain adequate positive and negative percent agreement for calling TMB high versus low across the cut-offs used in practice.2 Separate simulation work suggests panels between 1.5 and 3 Mb are best suited to estimating TMB with small confidence intervals, with smaller panels delivering imprecise estimates in the low-to-moderate range that matters clinically.5
Note what this figure does not say. It is not a claim that small panels are wrong. It is a claim that the same printed number means something different depending on the territory behind it, and the territory is not printed.
The counting rules, which are choices
Sampling variance sets a floor. Above it sit definitional choices, each defensible, each moving the value.
Germline filtering strategy
The gold standard for identifying somatic mutations is filtering against patient-matched germline sequencing. Most commercial panels do not sequence a matched normal, so germline variants are removed by reference to population databases instead — and that substitution inflates TMB, because private germline variants absent from the databases survive filtering and get counted as somatic.6
The magnitude is not small and it is not evenly distributed. In a cohort of 701 newly diagnosed multiple myeloma patients, TMB estimated with paired germline filtering was comparable between Black and White patients (6.09 vs 5.47 mut/Mb). Substituting public-database filtering inflated estimates for both groups, but significantly more for Black patients, with a race-by-filtering interaction p < 2×10⁻¹⁶.6 The mechanism is under-representation in the reference databases: the less well your ancestry is represented, the more of your germline variation looks novel, and the more of it is miscounted as tumor mutation.
Even within database filtering, the population minor allele frequency cut-off is a free parameter. Consortium work assessed pMAF thresholds of 0%, 0.5%, and 1%, finding that filtering potential germline variants at >0% pMAF gave the strongest correlation to WES TMB.2 Different laboratories set this differently, and the correlation between TMB values computed under different thresholds is high while the correlation between database-filtered and paired-germline-filtered values is considerably lower — the impact varies substantially from patient to patient.6
Synonymous variants
Some pipelines count only non-synonymous changes; others include synonymous ones. This is a documented, deliberate divergence — commercial tooling exposes it as two separate scores, one including synonymous variants and one excluding them.7
Its effect is measurable. In a comparison of exome-derived TMB, including synonymous variants raised the tumor-only value by about 5.76 mut/Mb relative to paired calling, versus about 3.88 mut/Mb when synonymous variants were excluded.8 The rationale for including them is not arbitrary — synonymous variants add signal to the mutation-rate estimate and reduce sampling noise, and one multicenter analysis concluded that including synonymous, nonsense, and hotspot mutations could enhance accuracy.3 Both choices are defensible. They are not interchangeable.
Pathogenic and hotspot variant handling
Panels are enriched for cancer genes, which is to say enriched for recurrently mutated positions. A panel's mutation density is therefore not representative of the genome's, and counting driver mutations at panel density and extrapolating to a per-megabase rate overestimates the genome-wide rate. The consortium found that failure to filter out pathogenic variants when estimating panel TMB resulted in overestimating TMB relative to WES for all assays.2 The bias is a direct consequence of what the panel was designed for.
The denominator
TMB is a ratio, and the denominator is not the panel's nominal size — it is the territory that actually achieved callable coverage in that specimen. Pipelines define this explicitly: eligible regions are the subset of the regions of interest that exceed a minimum coverage threshold.9 A degraded specimen with patchy coverage has a smaller real denominator than the panel's advertised footprint, which alters the rate even with the numerator unchanged. Two laboratories quoting the same nominal panel size may be dividing by different numbers.
The VAF floor
Which variants clear the reporting threshold depends on the minimum allele fraction, and that interacts with tumor purity. Recommendations exist — one multicenter study proposed a 5% VAF cut-off for samples with at least 20% tumor purity3 — but the choice is assay-specific, and, as covered previously in this series, must be raised for FFPE material to suppress deamination artifacts. A pipeline tuned for artifact suppression counts fewer variants than one tuned for subclonal sensitivity, on identical data.
| Parameter | Documented effect | Direction |
|---|---|---|
| Panel size | CV ~35% at TMB = 10 even for 1.1–1.4 Mb panels; >667 Kb needed for adequate agreement2,3 | Widens uncertainty; no directional bias, but misclassification rises near the threshold |
| Germline filtering | Database filtering vs paired normal inflates TMB, significantly more in under-represented groups6 | Inflates, unevenly across ancestries |
| Synonymous inclusion | ~5.76 vs ~3.88 mut/Mb tumor-only inflation with vs without synonymous variants8 | Raises the count; also lowers sampling noise |
| Pathogenic/hotspot filtering | Failure to filter pathogenic variants overestimated TMB relative to WES for all assays2 | Inflates, because panels are enriched for driver positions |
| Callable denominator | Eligible territory is the coverage-passing subset of target regions, not nominal panel size9 | Specimen-dependent; shrinks with degraded input |
| VAF floor | Interacts with purity; 5% proposed at ≥20% purity, higher floors used for FFPE3 | Higher floor lowers the count |
Figures are from the cited studies and reflect those cohorts, panels, and pipelines. The mechanisms generalize; the magnitudes are not constants.
MSI: the same problem with a different shape
MSI is not a rate but a fraction — the proportion of interrogated microsatellite loci showing instability — and it is thresholded. That structure creates its own comparability problem, and one detail makes it sharper than TMB's.
Panels interrogate different loci, in different numbers, and score them differently. FoundationOne CDx analyzes 95 intronic homopolymer repeat loci with adequate coverage and compiles them into an overall MSI score via principal components analysis; another assay in the same comparison used MSIsensor to report a percentage of unstable microsatellites, over a different gene set entirely.10 These are different measurements that share a name and a clinical category.
The decisive point is that the threshold and the locus list are not independent parameters. MSIsensor's recommended threshold is 3.5% of unstable sites, whereas mSINGS and MSI-PCR conventionally use 20%.11 That is not a disagreement about biology. The low MSIsensor threshold compensates for the inclusion of poorly performing loci in a large panel; apply that same threshold to a curated list of top-performing loci and specificity falls sharply. Conversely, the 20% threshold works well on a small panel of well-chosen loci but limits sensitivity when applied to a large one.11
A threshold is therefore only meaningful paired with the locus set it was calibrated on. An MSI score of 8% means "unstable" under one convention and "stable" under another, and neither the score nor the report carries the locus list that would disambiguate it.
Locus composition matters biologically too: mononucleotide repeats are more sensitive and specific for MSI than longer repeat motifs, and loci with longer motifs can simply inflate the denominator of the instability fraction.12 Two panels with the same nominal number of microsatellite loci but different motif composition are not equivalent instruments.
There is also a tumor-type dependency. In a 263-specimen comparison of PCR-based and NGS-based MSI testing, overall sensitivity and specificity of the NGS assay were 92.2% and 98.8% — but colorectal cases reached near-optimal concordance (98.1% sensitivity, 100% specificity) while endometrial cases showed only 88.6% sensitivity, driven by cases with instability at fewer than five markers.13 An assay validated in colorectal cancer does not automatically transfer to endometrial cancer, because the phenotype it is detecting is subtler there.
What is genuinely reassuring
This piece would be misleading if it stopped at the divergence, because the empirical agreement is, in the aggregate, good.
Panel-based TMB correlates strongly with WES-based TMB and is broadly comparable across gene panels in the 0 to 40 mut/Mb range, with observed variance traceable to gene composition, technical specifications, and pipeline.14 Calibration reduced spread around the WES value for 90% of clinical samples tested.2 For MSI, NGS-based detection is accurate — one comparison across 458 tumor-normal pairs and six cancer subtypes found all three tools tested classified samples correctly at high rates, with the best at 98.91% accuracy.11
So the practical situation is narrower than "these numbers are unreliable." It is this: agreement is strong in the middle of the range and degrades near the threshold, which is the only place the number is used for a decision. A tumor at 40 mut/Mb will be called high by any reasonable assay. A tumor at 30 will too. The disagreement concentrates at 8 to 12, where a patient's access to a therapy is determined, and where sampling variance on a small panel is at its widest relative to the distance from the cut-off.
That is also why calibration is the right response rather than despair. The consortium built and released a calibration tool precisely because the problem is tractable.1 Harmonization is the correct engineering answer to systematic, quantified divergence.
What this means for a pipeline
Report the definition alongside the number
A TMB is evaluable only as a tuple: value, panel and version, callable megabases achieved in this specimen, germline filtering strategy (paired normal, or database with which pMAF threshold), whether synonymous variants were counted, whether pathogenic and hotspot variants were filtered, and the VAF floor. "TMB 11" is not a result; "TMB 11 mut/Mb, non-synonymous only, 1.2 Mb callable, tumor-only with 0% pMAF database filtering" is. The same applies to MSI: the score, the number and type of loci, the caller, and the threshold with the locus set it was calibrated against.
Attach an interval, not just a point
Given the panel size and the observed count, a confidence interval is a few lines of arithmetic, and near a decision threshold it is the most informative thing on the page. A point estimate of 11 with an interval of 6 to 18 should not be actioned the way a point estimate of 11 with an interval of 10 to 12 is. Publicly available tools exist for exactly this visualization.4
Do not transfer a threshold between assays
This is the operational core. A cut-off is a property of an assay, validated on that assay's territory, counting rules, and locus set — not a property of the biomarker. Importing FoundationOne CDx's 10 mut/Mb onto a different panel without calibration means applying a decision boundary derived from a different measurement scale. The MSI case makes this vivid: 3.5% and 20% are both correct thresholds, for different instruments.
Calibrate rather than compare raw
Where cross-assay comparison is genuinely needed — cohort assembly, retrospective analysis, trial enrollment across sites — use the published calibration approaches rather than pooling raw values.1,2 And where a cohort mixes assays, treat assay identity as a covariate, not as a nuisance to be ignored. A biomarker whose measurement scale varies by site, entering an association analysis unmodelled, is the same structure this series has described in the population stratification and batch confounding pieces.
Prefer paired-normal filtering where the question is quantitative
Tumor-only germline filtering is a practical necessity in many settings, but it introduces an inflation that is patient-specific and ancestry-dependent.6 Where TMB is being used near a threshold, and particularly in diverse patient populations, the case for a matched normal is a quantitative one, not merely best practice.
The shape of the problem
There is a recurring structure in this series: a pipeline performs a resolution that is invisible in its output. Protein inference resolves peptide ambiguity by parsimony and reports an accession. SV callers resolve breakpoint placement within a tolerance window and report a coordinate. Deconvolution resolves missing cell types by redistribution and reports a proportion.
TMB and MSI are the same structure at the level of an entire assay definition. A laboratory makes a dozen reasonable methodological choices — territory, filtering strategy, variant classes, thresholds, denominators — and the output is one number compared against one cut-off. The choices are documented in a validation package that the treating oncologist will never see, and the number arrives on the report wearing the same clothes regardless of how it was made.
The field knows this. The consortium quantified it across 16 laboratories, built the calibration tool, and published the recommendations. What has not changed is the interface: a scalar and a threshold, with the definition left behind in the methods.
References
- Friends of Cancer Research. Tumor Mutational Burden (TMB) Harmonization Project. https://friendsofcancerresearch.org/tmb/
- Vega DM, Yee LM, McShane LM, et al. Aligning tumor mutational burden (TMB) quantification across diagnostic platforms: phase II of the Friends of Cancer Research TMB Harmonization Project. Annals of Oncology. 2021;32(12):1626–1636. https://pubmed.ncbi.nlm.nih.gov/34606929/
- Budczies J, Allgäuer M, Litchfield K, et al. Optimizing panel-based tumor mutational burden (TMB) measurement. Annals of Oncology. 2019;30(9):1496–1506. https://www.annalsofoncology.org/article/S0923753419459930/pdf
- Visualization of the effect of assay size on the error profile of tumor mutational burden measurement. Genes. 2022;13(3):432. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8949329/
- Buchhalter I, Rempel E, Endris V, et al. Size matters: dissecting key parameters for panel-based tumor mutational burden analysis. International Journal of Cancer. 2019;144(4):848–858. https://www.researchgate.net/publication/327819355_Size_matters_Dissecting_key_parameters_for_panel-based_tumor_mutational_burden_analysis
- Inflation of tumor mutation burden by tumor-only sequencing in under-represented groups. npj Precision Oncology. 2021;5:22. https://www.nature.com/articles/s41698-021-00164-5
- VarSome Clinical documentation. Tumor mutational burden estimation. https://docs.varsome.com/en/tmb (vendor documentation, cited for the synonymous/non-synonymous scoring distinction)
- Analysis of tumor mutational burden: correlation of five large gene panels with whole exome sequencing. Molecular Cancer / BMC. 2020. https://pmc.ncbi.nlm.nih.gov/articles/PMC7347536/
- Illumina DRAGEN documentation. Tumor mutational burden. https://support-docs.illumina.com/SW/dragen_v42/Content/SW/DRAGEN/Biomarkers_TMB.htm (vendor documentation, cited for the callable-region definition)
- Concordance analysis of microsatellite instability status between polymerase chain reaction based testing and next generation sequencing for solid tumors. Scientific Reports. 2021;11:19934. https://www.nature.com/articles/s41598-021-99364-z
- Kautto EA, Bonneville R, Miya J, et al. Performance evaluation for rapid detection of pan-cancer microsatellite instability with MANTIS. Oncotarget. 2017;8(5):7452–7463. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5352334/
- A robust microsatellite instability detection model for unpaired colorectal cancer tissue samples. Chinese Medical Journal. 2022. https://mednexus.org/doi/10.1097/CM9.0000000000002216
- Concordance in detection of microsatellite instability by PCR and NGS in routinely processed tumor specimens of several cancer types. Cancer Medicine. 2023. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10501280/
- Fancello L, Gandini S, Pelicci PG, Mazzarella L. Tumor mutational burden quantification from targeted gene panels: major advancements and challenges. Journal for ImmunoTherapy of Cancer. 2019;7:183. https://jitc.biomedcentral.com/articles/10.1186/s40425-019-0647-4

