Neoantigen Prediction

Neoantigen Prediction: Why One Label Carries Three Different Claims
Zetobit · Bioinformatics Insight Series Immuno-oncology
BIOINFORMATICS INSIGHT SERIES Neoantigen Prediction Why One Label Carries Three Different Claims ONE WORD, THREE QUESTIONS Does it bind? chemistry — predicted well Is it presented? cell biology — partly Is it recognised? immunology — barely Across 25 independent pipelines given identical data, 6% of top-ranked peptides were immunogenic. Kanna Nandakumar, PhD ZETOBIT
Bioinformatics Insight Series

Neoantigen Prediction

Why One Label Carries Three Different Claims

A neoantigen pipeline takes tumour and normal sequencing, finds somatic mutations, translates the ones in expressed genes, slides a window along the mutant protein, and scores each resulting peptide against the patient's HLA alleles. The output column says predicted neoantigens. That single label compresses three claims that rest on different evidence and are true at very different rates.

The peptide binds the MHC molecule. The peptide is actually processed and displayed on the cell surface. And a T cell in that patient's repertoire recognises it and responds. Only the third makes a peptide a therapeutic target, and it is the one the tools know least about.

The consortium experiment

The cleanest evidence comes from an exercise designed to remove the usual confound of everyone using different data. The Tumor Neoantigen Selection Alliance gave whole-exome and RNA sequencing from six tumours — three melanoma, three non-small-cell lung cancer — to teams from academia, pharma and biotech, and asked each to submit ranked neoantigen predictions1.

Twenty-eight teams submitted; 25 were analysed. Team submissions ranged from 7 to 81,904 ranked peptide-MHC pairs per tumour, with a median of 204. From the top-ranked peptides across all groups, 608 were tested for immunogenicity by multimer-based assays, and 37 — 6% — were immunogenic. Each individual team had a median of 51 of its peptides tested, of which a median of 3 (again 6%) were immunogenic2.

Two things about that rate matter. It matched what had previously been reported2, so it was not an artefact of this particular exercise. And because every team landed at the same 6%, the number describes the difficulty of the question rather than the quality of any one pipeline.

Where the pipelines disagreed — and where they did not

Given identical input data, the predictions diverged sharply. Analysis of the top 100 predicted peptides from each group revealed a median overlap of 13%3, and the limited overlap could not be explained simply by differences in variant calling4.

So the divergence was not upstream. Everyone was looking at broadly the same mutations.

Now the finding that reframes the whole problem. Of the 25 teams analysed, 20 submitted predicted MHC binding — and there were no substantial differences between teams in how well they predicted MHC binding strength5.

The pipelines agreed almost completely on the chemistry and almost not at all on the ranking. The disagreement sits precisely where the tools have the least evidence, not where they have the most.

This is worth dwelling on, because it inverts the intuitive reading. A 13% overlap between two pipelines looks like a field where nothing works. It is closer to a field where one layer works well, everyone implements that layer competently, and the layers above it — which peptides to prioritise, and on what grounds — are where the choices live and where the evidence runs out.

Binding prediction is the part that works

Being fair to the tools here matters, because the argument is not that MHC binding predictors are unreliable.

They are among the better-validated predictors in computational biology. For many years they were trained primarily on quantitative peptide-HLA binding affinity data from in vitro experiments; more recent methods are trained on mass-spectrometry immunopeptidomics data — thousands of peptides observed actually bound to MHC in patient samples and cell lines6. That shift matters, because eluted-ligand data folds processing and presentation into the training target rather than pure binding chemistry.

And binding really is necessary. One analysis of natural T-cell responses found that immunogenic neopeptides were predicted to bind significantly more strongly to HLA than non-immunogenic ones, with 96% of the immunogenic peptides sharing very strong predicted binding, comparable to pathogen-derived epitopes7.

Necessary, and nowhere near sufficient. A tumour may contain hundreds of candidate altered peptides, of which few will bind MHC and elicit a response6 — and the filter that removes non-binders leaves a set still dominated by binders that go nowhere.

What sits between binding and presentation

The gap between the first and second claims is not a modelling subtlety; it is a sequence of physical steps a peptide has to survive.

The mutant protein has to be made, which requires the transcript to be expressed in that tumour at that time — and expression filters use bulk RNA-seq, so a peptide from a gene expressed in a subclone can pass a whole-tumour threshold it would fail in most cells. The protein has to be degraded by the proteasome with cuts falling in the right places to liberate that exact peptide. The fragment has to be transported into the endoplasmic reticulum by TAP, which has its own sequence preferences. It has to be loaded onto an MHC molecule that the cell is actually expressing — and HLA loss, both allele-specific and total, is a recognised immune-escape mechanism, so the allele a prediction was made against may be absent from the cells that matter.

Each step is a filter, and none of them is fully captured by an affinity score. This is why the shift to immunopeptidomics training data was substantive rather than incremental: eluted-ligand models learn the composite outcome of all of these steps at once, because the training examples are peptides that survived every one of them6. What those models cannot learn from is anything about the patient's T-cell repertoire, which is not in the data at all.

The immunogenicity layer

Tools that predict immunogenicity directly exist. Their measured performance is the sobering part.

A curated benchmark of 199 tumour-specific neoantigens with experimentally validated MHC-I presentation and known positive or negative immune response evaluated sixteen metrics as immunogenicity predictors. At the classification thresholds used, the best performers were a MixMHCpred score at 30% false positives and a NetMHCpan rank at 31%; the worst were DeepHLApan at 100% and DeepImmuno at 79%8.

A 100% false-positive rate on a curated set is not a small calibration issue. And note which tools posted the two best figures: general-purpose binding predictors, not the dedicated immunogenicity models.

The errors also run the other way, which is easy to forget when the discussion is about false positives. In one comparison against a functional identification pipeline, 12 of 49 peptides (25%) with confirmed IFN-γ responses would not have been identified by conventional in silico prediction, because their predicted binding affinity was above the 500 nM threshold9. A quarter of the confirmed responses sat outside the standard binding cut-off.

Why the third claim is a different kind of problem

It would be easy to read the immunogenicity numbers as a field that simply needs better models. Part of it is that. But the third claim differs from the first two in a way that no amount of training data fully resolves.

Binding is a property of two molecules. Presentation is a property of a cell. Recognition is a property of a particular patient's T-cell repertoire — which is shaped by thymic selection against that individual's own proteome, by their infection history, and by whatever clonal expansions and deletions have already happened. Two patients with identical HLA types and the identical mutation can differ in whether a T cell capable of seeing it exists at all.

A model trained on pooled data across patients is estimating an average propensity over repertoires. That is a real and useful quantity, and it is not the same as the question being asked, which is about one person. The self-similarity features that improve these models — peptide dissimilarity to self was found to predict immunogenicity where neo- and normal peptides had comparable predicted binding7 — are essentially proxies for how likely a repertoire is to have escaped tolerance, which is exactly the right instinct and still a population-level answer to an individual-level question.

The benchmarks have a circularity problem of their own

There is a structural issue with how these tools get evaluated, and it is the same shape as the one covered in the piece on variant-effect predictor benchmarks — arriving here from a different direction.

Validated neoantigen datasets are almost always assembled from peptides that were selected for testing because a binding predictor ranked them highly. One review of immunogenicity tools makes the consequence explicit: in datasets composed of peptides preselected by a criterion such as predicted MHC binding, that criterion will generally show no predictive performance, and the bias extends to every model that directly or indirectly predicts binding and presentation. The authors excluded NetMHCpan from their evaluation for exactly this reason, because their own dataset had been selected using it10.

The consequence is uncomfortable in both directions. Binding predictors look artificially weak on these benchmarks, because the range of binding strengths has been truncated by the selection. And a peptide that is genuinely immunogenic but a poor predicted binder is unlikely to be in any validated dataset at all — so the false-negative rate of the whole approach is systematically under-observed.

25 PIPELINES, THE SAME SIX TUMOURS Where they agree No substantial differences between teams in how well they predicted MHC binding strength. The chemistry layer is not the disputed one. Where they diverge Median overlap between any two teams' top 100 peptides: 13%. Submissions ranged from 7 to 81,904 ranked peptides per tumour. The limited overlap was not explained by differences in variant calling. The disagreement sits where the tools have the least evidence — not where they have the most. WHAT SURVIVED VALIDATION 37 of 608 top-ranked peptides were immunogenic — 6% Each individual team fared the same: a median of 51 of its peptides tested, a median of 3 confirmed. The rate is a property of the question rather than of any one pipeline — which is why picking a better tool is not the available remedy.
Figure 1. The consortium result. The 13% overlap figure is from a commentary summarising the same study; the binding-agreement statement is from the consortium's own presented analysis of the 20 teams that submitted binding predictions.
Three claims inside one label
ClaimWhat it rests onWhat it does not establish
The peptide binds MHC Affinity or eluted-ligand models, trained on large experimental datasets6 That the peptide is ever produced in the cell
The peptide is presented Immunopeptidomics-trained models; expression and clonality filters That any T cell exists which can see it
A T cell responds Immunogenicity models with 30–100% false-positive rates on curated sets8 That the response is useful in the tumour microenvironment
The candidate list as a whole The composition of all three layers Roughly 6% validated in the consortium test2

The fourth row is not the product of the three rows above it — it is what was observed when top-ranked candidates from real pipelines were tested. Note also that even a confirmed T-cell response in an assay is a fourth claim short of clinical benefit, which depends on trafficking, tumour heterogeneity and the local immune environment.

What to do about it

  1. Name the claim in the column header. “Predicted binders,” “predicted presented peptides” and “predicted immunogenic peptides” are three different lists; calling all three “neoantigens” is where the inference silently happens.
  2. Expect roughly one in twenty, and design around it. A 6% validation rate is the field's observed baseline2, not a sign of a bad pipeline. Peptide synthesis and assay capacity should be planned against that rate.
  3. Do not treat pipeline agreement as strong corroboration, or disagreement as failure. With a 13% median overlap on identical data3, two pipelines converging on a peptide is genuinely informative — and one pipeline missing it says little.
  4. Keep the binding layer and the ranking layer separate in the report. They have different evidence bases and different failure rates, and the consortium found the teams differed on the second while agreeing on the first5.
  5. Ask what any published immunogenicity benchmark was selected on. If candidates were chosen by a binding predictor, the benchmark cannot fairly evaluate binding-based methods10 — and it under-observes the false negatives entirely.
  6. Do not use a hard binding threshold as the only gate where sensitivity matters: a quarter of confirmed responses in one comparison fell outside 500 nM9.
  7. Record HLA typing quality alongside the predictions. Every score is conditional on the alleles supplied; a mistyped allele changes the entire candidate list without changing anything visible in the output.
  8. Report the validated fraction, not just the candidate count. A pipeline that outputs 81,904 ranked peptides and one that outputs 7 are not comparable objects2, and only validation distinguishes them.
  9. Treat a confirmed T-cell response as evidence about recognition, not about efficacy. Everything between recognition and tumour control is outside what the pipeline modelled.

The shape of the error

This series keeps finding the same structure: a computation that is correct, carrying an inference it does not support. The neoantigen case has a particularly clean version, because the compression happens in a word rather than in a number.

“Neoantigen” is an immunological term. It means a mutant peptide that the immune system actually responds to. What a pipeline produces is a peptide that scored well on a model of one necessary step, filtered by expression and a few heuristics. Applying the immunological word to the computational output moves a claim about chemistry into a claim about immunity, and nothing in the file records that the move was made.

What makes this more tractable than most entries in the series is that the field measured its own attrition, publicly, with 25 pipelines on shared data and a central validation experiment. The 6% is not a criticism from outside; it is the consortium's own headline. And knowing the rate is what makes a candidate list usable — you plan for twenty peptides to find one, rather than treating the list as twenty targets.

The mutations are real. The binding predictions are good. What the label promises is a T-cell response, and that is the one thing in the chain that no model in the pipeline has seen.

References

  1. Wells DK, Dang KK, Hubbard-Lucey VM, et al. Key parameters of tumor epitope immunogenicity revealed through a consortium approach improve neoantigen prediction. Cell 2020;183(3):818–834. cell.com — S0092-8674(20)31156-9
  2. Wells DK, et al., as ref. 1 — team submission ranges, the 608 tested peptides, and the 37 (6%) immunogenic result. sciencedirect.com — S0092867420311569
  3. Strength in numbers: identifying neoantigen targets for cancer immunotherapy. Cell 2020 (commentary on ref. 1); source of the 13% median overlap figure for top-100 predicted peptides. Author list not captured — verify before publication. cell.com — S0092-8674(20)31317-9
  4. Abstract 3210: Strategies to improve the sensitivity and ranking ability of neoantigen prediction methods — report on TESLA results. Cancer Research 2020;80(16 Suppl):3210. Conference abstract; cited for the statement that limited overlap was not explained by variant-calling differences. aacrjournals.org — Cancer Res 80(16 Suppl):3210
  5. Elucidation of key parameters for effective neoantigen prediction through a consortium effort: the Tumor nEoantigen SeLection Alliance. Presented slide deck accompanying the TESLA study; source of the statement that 20 of 25 teams submitted binding predictions with no substantial between-team differences. NOT a peer-reviewed venue — confirm against the published paper before quoting. TESLA consortium presentation (PDF)
  6. Sarkizova S, et al. / MHCnuggets: high-throughput prediction of MHC class I and class II neoantigens. bioRxiv 752469. Preprint; cited for training-data provenance (in vitro affinity vs immunopeptidomics) and the candidate-attrition framing. Author list not fully captured — verify. biorxiv.org — 752469
  7. An analysis of natural T-cell responses to predicted tumor neoepitopes. Frontiers in Immunology 2017. Author list not captured — verify before publication. ncbi.nlm.nih.gov/pmc/articles/PMC5694748
  8. Unraveling tumor specific neoantigen immunogenicity prediction: a comprehensive analysis. Frontiers in Immunology 2023;14:1094236. ITSNdb benchmark of 199 tumour-specific neoantigens. Author list not captured — verify before publication. frontiersin.org — fimmu.2023.1094236
  9. Comparison of a functional neoantigen identification pipeline against conventional in silico prediction, reporting that 12 of 49 IFN-γ-positive peptides fell above the 500 nM binding threshold. Source is a granted US patent specification rather than a peer-reviewed paper — WEAKEST reference here; replace with a primary publication before use. USPTO 12427195
  10. Beyond MHC binding: immunogenicity prediction tools to refine neoantigen selection in cancer patients. Exploration of Immunology 2023. Author list not captured — verify before publication. explorationpub.com — ei/100391
Zetobit, LLC · Bioinformatics consulting · Lexington, KY · zetobit.com
Previous
Previous

Reading a VUS

Next
Next

Epigenetic Clocks