Immune Repertoire Sequencing: Why a Clone Frequency Is an Amplification Artifact and a Clone Is a Threshold You Chose

Immune Repertoire Sequencing: Why a Clone Frequency Is an Amplification Artifact and a Clone Is a Threshold You Chose — Zetobit
Zetobit.
BIOINFORMATICS INSIGHT SERIES Immune Repertoire Sequencing Why a clone frequency is an amplification artifact, and a clone is a threshold you chose EQUAL INPUT equimolar spike-ins multiplex PCR OBSERVED OUTPUT systematic, reproducible distortion The bars moved. No cell expanded. ZETOBIT. Kanna Nandakumar, PhD

Bioinformatics Insight Series

Immune Repertoire Sequencing: Why a Clone Frequency Is an Amplification Artifact and a Clone Is a Threshold You Chose

TCR and BCR sequencing produce two numbers that look like counts — how many clones are present, and how large each one is. Both are reconstructions. One is distorted by the chemistry that made the library; the other depends on a clustering parameter set by the analyst.

Repertoire sequencing has an unusually direct clinical story. A T cell that recognizes a tumor antigen proliferates; its receptor sequence becomes more abundant; sequencing the receptors should show that clone rising. Clonal expansion is one of the few places in genomics where the measurement seems to map onto the biology almost without translation.

The translation is the problem. Between a T cell in a tumor and a frequency in a report sit a PCR reaction with dozens of competing primers, a sampling step that sees a small fraction of the repertoire, and a clustering decision about which sequences count as the same clone. Each of those transforms the number. None of them is visible in the output, which arrives as a clean ranked table of clonotypes and percentages.

Multiplex PCR does not amplify all receptors equally

The V(D)J diversity that makes the repertoire informative is also what makes it hard to amplify. A bulk assay must capture rearrangements spanning many V and J segments, which in the standard DNA-based approach means a multiplex reaction with a large primer pool — one protocol uses 20 V-specific forward primers and 13 J-specific reverse primers in a single tube.

Those primers do not perform identically, and the resulting distortion is large. Using synthetic antibody spike-in genes to isolate the effect, one study found that primer bias from multiplex PCR library preparation resulted in antibody frequencies with only 42% to 62% accuracy. Comparing spike-in frequencies from singleplex versus multiplex amplification of the same material produced a correlation of R² = 0.56 — meaning nearly half the variance in observed frequency was contributed by the amplification step rather than by the input.

Two properties of that bias matter more than its size.

First, it is systematic, not random. The same authors observed that variation across replicates was extremely low and spike-ins sharing the same V-gene were consistently under- or over-amplified. Replicate concordance therefore proves nothing about accuracy — running the assay twice reproduces the same distortion, which is exactly what makes it easy to mistake for signal.

Second, it is specific to the primer set. Comparing uncorrected spike-in frequencies generated with two different multiplex primer sets yielded R² = 0.08. Two laboratories measuring the same sample with different primer pools would produce frequency profiles with almost no relationship to each other. After correction with molecular amplification fingerprinting, the same comparison rose to R² = 0.84.

Reproducibility and accuracy come apart here. The bias is stable across replicates, which means running it again confirms the artifact rather than exposing it.

The correction strategies all amount to measuring the bias rather than assuming it away. Unique molecular identifiers tag individual starting molecules before amplification so that final read counts can be collapsed back to input molecules — the MAF approach applied UID tagging before and during multiplex PCR and recovered frequencies with up to 99% accuracy. Synthetic template spike-ins take the complementary route: one protocol adds 260 synthetic TCR templates at equimolar concentration as internal controls, so per-primer efficiencies can be measured directly and scaled out. RNA-based 5'RACE methods sidestep much of the problem by needing only a single primer pair, at the cost of a different efficiency penalty at the ligation step.

Diversity is worse than frequency

If frequencies are distorted, diversity metrics are distorted more, because they are dominated by the rare clones that are hardest to measure.

Sequencing error is the first amplifier. The same spike-in study found that Ig-seq errors resulted in antibody diversity measurements being overestimated by up to 5000-fold. Every PCR or sequencing error in a highly abundant clone creates a sequence that has never existed in a cell and that a naive pipeline counts as a new, rare clonotype. Because errors scale with read depth, sequencing deeper generates more phantom clones.

Sampling is the second. Repertoire diversity estimation is a version of what ecologists call the missing-species problem: a clinical blood draw contains a small fraction of an overall repertoire that may comprise billions of cells, so some clones — especially small ones — almost always go unsampled and undetected, and sample diversity usually underestimates true diversity. Reviews of the classical richness estimators applied to TCR data conclude that existing approaches have significant shortcomings and frequently underestimate true diversity, and note that PCR amplification can overestimate repeated observation of a clonotype, producing false saturation and further underestimates.

The two errors push in opposite directions, which is the awkward part. Error inflates apparent diversity; undersampling deflates it. A reported diversity index is the net of two large opposing biases whose magnitudes depend on read depth, input cell number, error rate, and clustering stringency. Comparing that index across samples processed differently is not a comparison of immune systems.

SEVEN RELATED BCR SEQUENCES, VARYING BY SOMATIC HYPERMUTATION pairwise junction distances range from 2% to 14% Threshold 0.20 1 clone one large expanded lineage Threshold 0.10 3 clones three moderate lineages Threshold 0.02 7 clones no expansion detected — high diversity

Figure 1. For B cells, "the same clone" is not given by the data. Somatic hypermutation means clonally related receptors differ from one another, so partitioning requires a distance threshold. The same seven sequences yield one expanded lineage, three moderate ones, or seven singletons depending on where the cut is placed — and clonality and diversity metrics move with it. Distances shown are illustrative.

For B cells, "the same clone" is a parameter

T cell receptors do not hypermutate, so identical TCR sequences can reasonably be grouped as one clonotype. B cells break that convenience. During affinity maturation, activated B cells expand and undergo somatic hypermutation of their receptor, forming a clone of diversified cells that can be related back to a common ancestor. Members of one clone are, by construction, not identical.

Clonal grouping therefore requires deciding how different two sequences can be while still counting as the same lineage. The standard approach partitions sequences sharing the same IGHV gene, IGHJ gene, and junction length, then applies hierarchical clustering on junction distance — with the threshold chosen by finding the valley in a bimodal distance-to-nearest distribution, since one mode reflects clonally related sequences and the other unrelated ones. Automated methods fit that bimodal distribution explicitly with a mixture model to place the cut and to estimate study-specific sensitivity and specificity.

This works well when the distribution is cleanly bimodal. It is nonetheless a choice, and one worth stating in a methods section, because the number of clones and the size of the largest clone both move with it. Newer methods use adaptive rather than fixed thresholds, or add information beyond junction similarity — one approach combines junction distance with shared somatic hypermutation events in the V and J segments, on the reasoning that clonally related sequences inherit mutations from a common ancestor.

One detail from the threshold literature is worth carrying into practice, because it inverts a natural instinct. Sequencing errors inflate the measured distance between genuinely clonal sequences — and the recommendation is to accommodate that rather than correct for it. If members of a clone were truly less than 10% different but experimental error increased their difference to under 11%, the proper choice is to use 11% as the threshold. The threshold is calibrated to the data as it exists after error, not to the biology as it would be without it.

Table 1. Four transformations between a cell in tissue and a number in a repertoire report. Each is a modeling step; none appears in the output table.
Step What it changes Control
Sampling Which clones are present at all; rare clones missed Report input cell number; treat richness as a lower bound
Multiplex amplification Relative frequencies, systematically and reproducibly UMIs before amplification; synthetic spike-in standards
Sequencing error Inflates apparent diversity, creates phantom rare clones UMI consensus; error-aware clustering
Clonal grouping How many clones exist and how large the largest is Data-driven threshold, reported with the result

Where this lands clinically

Repertoire metrics are increasingly used as immunotherapy biomarkers. Reviews describe intratumoral clonality and peripheral diversity as correlating with outcomes — a focused, clonal intratumoral repertoire often associated with improved survival, high baseline tumor clonality frequently correlating with response to PD-1/PD-L1 inhibitors, and greater peripheral diversity potentially predicting benefit from anti-CTLA-4 therapy.

Both of those quantities are exactly the ones the technical pipeline distorts. Clonality is a function of the frequency distribution, which multiplex PCR reshapes. Diversity is a function of rare-clone detection, which error inflates and sampling deflates. A biomarker built on either is a biomarker built on an assay-specific reconstruction.

That does not invalidate the biology. It does mean the biomarker travels only as far as the protocol does. A clonality threshold established with one primer set, input amount, and clustering parameter has no defined meaning applied to data generated differently — and because the distortions are systematic, the mismatch produces a consistent, confident, wrong answer rather than obvious noise.

What to fix in the pipeline

Put UMIs upstream of amplification where the chemistry allows. Collapsing reads to input molecules addresses frequency bias and error inflation simultaneously. This is straightforward for RNA-based protocols and remains harder for DNA-based multiplex approaches, where efficient incorporation before PCR is still challenging.

Run spike-in standards when UMIs are not available. Synthetic templates at known concentration measure per-primer efficiency directly. Because bias parameters are largely preserved across datasets using the same primer pool, scaling factors derived from a subset of samples can be applied per batch rather than run on every sample.

Report the clonal-grouping threshold and how it was chosen. A clonality figure without its threshold is not interpretable and not reproducible. Where the distance-to-nearest distribution is not cleanly bimodal, that is a QC finding, not a detail.

Fix the protocol across a longitudinal series. Because bias is primer-set-specific, changing chemistry mid-study converts a technical change into an apparent immunological one. This is the same constraint that applies to basecaller models and batch effects, with the same remedy: change everything or nothing.

Treat richness as a lower bound. Report sampling depth and input cell counts alongside any diversity metric, and prefer measures less sensitive to missing rare species when comparing across samples — while noting that weighted measures answer a different question than richness does.

The short version

A repertoire report presents clone frequencies and diversity as if they were counted. They are reconstructed through sampling, a biased amplification, an error process, and a clustering threshold. The amplification bias is systematic and primer-set-specific, so replicates agree while measurements from different protocols do not — and the clone itself, for B cells, is defined by a parameter rather than found in the data.

A measurement that looks like counting

Most of the failure modes in this series involve a model output presenting as an observation. Repertoire sequencing has a sharper version of the problem, because the underlying quantity really is a count — there is a true number of cells carrying each receptor, and it is an integer. That makes the output feel like arithmetic rather than inference.

It isn't. The path from cells to percentages runs through a chemistry that favors some sequences over others, a sample that misses most of the rare ones, an error process that invents new ones, and a threshold that decides which are the same. The number at the end is real and useful. It is just not a count, and treating it as one is how a change of primer pool becomes a change in the immune system.

Zetobit builds and validates CAP/CLIA-compliant NGS pipelines, including AIRR-seq workflows with UMI-based consensus, spike-in bias correction, and documented clonal-grouping thresholds. If you are standing up TCR/BCR profiling for an immuno-oncology program or reconciling repertoire metrics across protocols, we're happy to talk.

References

  1. Khan TA, Friedensohn S, de Vries ARG, et al. Accurate and predictive antibody repertoire profiling by molecular amplification fingerprinting. Science Advances 2(3):e1501371 (2016). doi:10.1126/sciadv.1501371.
  2. Gupta NT, Adams KD, Briggs AW, et al. Hierarchical clustering can identify B cell clones with high confidence in Ig repertoire sequencing data. Journal of Immunology 198(6):2489–2499 (2017). doi:10.4049/jimmunol.1601850.
  3. Nouri N, Kleinstein SH. Optimized threshold inference for partitioning of clones from high-throughput B cell repertoire sequencing data. Frontiers in Immunology 9:1687 (2018). doi:10.3389/fimmu.2018.01687.
  4. Nouri N, Kleinstein SH. Somatic hypermutation analysis for improved identification of B cell clonal families from next-generation sequencing data (SCOPer). PLOS Computational Biology 16(6):e1007977 (2020). doi:10.1371/journal.pcbi.1007977.
  5. Kaplinsky J, Arnaout R. Robust estimates of overall immune-repertoire diversity from high-throughput measurements on samples. Nature Communications 7:11881 (2016). doi:10.1038/ncomms11881.
  6. Laydon DJ, Bangham CRM, Asquith B. Estimating T-cell repertoire diversity: limitations of classical estimators and a new approach. Philosophical Transactions of the Royal Society B 370:20140291 (2015). doi:10.1098/rstb.2014.0291.
  7. An open protocol for modeling T cell clonotype repertoires using TCRβ CDR3 sequences. BMC Genomics 24 (2023). doi:10.1186/s12864-023-09424-z.
  8. Harnessing TCR repertoires: predictive insights and therapeutic monitoring in cancer immunotherapy (2025).
Previous
Previous

Spatial Deconvolution: Why a Cell Type Map Shows You the Reference, Not the Tissue

Next
Next

Tumor Purity and Ploidy: Why the Number Every Somatic Result Depends On Has More Than One Valid Answer