Splice-Effect Predictors: Why a 0.9 Is a Probability, Not a Consequence
Splice-Effect Predictors: Why a 0.9 Is a Probability, Not a Consequence
A SpliceAI delta score of 0.9 is often read as near-certainty that a variant is pathogenic through a splicing mechanism. It is not that. The score is an estimate of the probability that splicing is altered — a claim about whether the transcript changes, not about what the changed transcript is, how much of the mRNA pool it represents, or whether the resulting protein is broken. Those are three separate questions, and the number answers none of them. The threshold at which the score becomes evidence, meanwhile, is not a property of the tool: it varies by variant class, by gene, and by whether you are arguing for pathogenicity or against it.
Splice prediction is one of the genuine successes of deep learning in clinical genomics. Before it, deep intronic variants were essentially uninterpretable, and the tools available for splice-region variants were position weight matrices that performed poorly outside the canonical dinucleotides. SpliceAI changed what is findable. Nothing in this piece argues otherwise.
The problem is downstream of the model. A delta score is a continuous output that gets compressed, at some point in every pipeline, into a binary — flagged or not, evidence or not — and the compression discards information the interpretation depends on. Worse, the number looks like it means something specific, because it is bounded between 0 and 1 and gets called a probability. It is a probability of the wrong thing for the purpose it is usually put to.
This series covered the clinical payoff of pairing WGS with RNA-seq: resolving deep intronic and synonymous splice variants that DNA sequencing is uniquely good at finding and uniquely bad at interpreting. This piece is the other half of that argument. It is about why the prediction, on its own, does not close the case — and what specifically it leaves open.
What the score actually estimates
SpliceAI outputs four delta scores per variant — acceptor gain, acceptor loss, donor gain, donor loss — and the commonly used value is the maximum of the four. The tool's own documentation defines it plainly: the delta score ranges from 0 to 1 and can be interpreted as the probability of the variant being splice-altering.1
Read that definition against what a variant curator needs. "Splice-altering" means the splicing pattern differs from reference. It does not specify which of several possible aberrations occurs, what proportion of transcripts carry it, whether the reading frame survives, or whether the product is degraded by nonsense-mediated decay. A variant that causes 5% exon skipping and one that abolishes an exon entirely can both be splice-altering.
Tooling documentation states this limitation directly: SpliceAI does not predict the functional consequence of aberrant splicing — exon skipping, intron retention, and so on — only the probability that splicing is disrupted.2 The score is a detection statistic. Clinical interpretation needs a characterization.
This is the same structure the series described for structural variants, where detection, placement, and genotype fail independently and only detection is captured by the headline metric. Here the composite claim is smaller but the compression is identical: one number covering a question it was built for, and several it was not.
The threshold is not a property of the tool
If the score were a calibrated probability of pathogenicity, a single cut-off would serve. It is not, and it does not. The published thresholds diverge by more than an order of magnitude, and each is correct for its purpose.
The developers characterized three operating points rather than endorsing one: 0.2 for high recall, 0.5 as the general recommendation, and 0.8 for high precision.1 That is an appropriate way to present a detection tool, and it already signals that the choice belongs to the user.
Clinical variant curation landed considerably lower. The ClinGen Sequence Variant Interpretation Splicing Subgroup calibrated the score against ACMG/AMP evidence codes and set PP3 at a maximum delta score of 0.2 or above and BP4 at 0.1 or below, with BP7 combinable with BP4 for synonymous and intronic variants.3,4 The optimal BP4 threshold of ≤0.1 gave 87% specificity for variants outside the donor/acceptor ±1,2 dinucleotides but only 73% specificity for variants within the splice region.4 Same threshold, same tool, different specificity depending on where the variant sits.
Gene-specific expert panels calibrate differently again. The ClinGen Myeloid Malignancy Expert Panel adopted SpliceAI as its primary splicing predictor for RUNX1 with PP3 at ≥0.38 and BP4 at ≤0.20, thresholds established at 90% sensitivity and 90% specificity respectively in that gene's context.5 A variant scoring 0.3 is supporting evidence for pathogenicity under the general SVI recommendation and is not under the RUNX1 specification.
Independent gene-level evaluations produce their own optima. In a study of 285 NF1 variants, ROC analysis put the optimal cut-off at >0.22, with AUC 0.975 and sensitivity and specificity both around 94%.6 A calibration against curated minigene results found that a score of 0.285 gave 90% sensitivity and 90% specificity for separating no-or-low from intermediate-to-complete aberration.7
And in the other direction entirely: for detecting spliceogenic intronic variants more than 50 bp from an exon, sensitivity at the 0.5 threshold was originally reported at just 41%, while lowering the threshold to 0.05 for variants beyond 20 bp raised observed sensitivity to 94%.8 A cut-off that is sensible for splice-region variants misses most deep intronic events.
The pattern is worth naming. A benign-evidence threshold and a pathogenic-evidence threshold are not two ends of one scale with a gap between them; they are two separate calibrations, because the cost of a false call differs in each direction. That is why the SVI recommendation leaves a deliberate indeterminate band between 0.1 and 0.2 where neither code applies. Scores land in that band routinely, and a pipeline that forces a binary at 0.5 silently converts "no evidence either way" into "no evidence of splicing."
Where the calibration comes from, and why it does not transfer
The reported performance of splice predictors depends heavily on what they were evaluated against, and this is the mechanism behind the divergence above.
SpliceAI's headline performance — an area under the precision-recall curve of 0.98 — was computed on RNA-seq data.9 When the same tool was benchmarked on clinically relevant variant sets using functional splice assays, PR-AUC fell to 0.93 for ABCA4 non-canonical splice site variants, 0.91 for ABCA4 deep intronic variants, and 0.74 for MYBPC3 non-canonical splice site variants.9 The authors note that performance on clinically relevant variants varies considerably from the original paper, and attribute the difference to the RNA-seq evaluation set.9
Three things follow. Performance is gene-dependent, in a way a genome-wide threshold cannot accommodate. Performance is variant-class dependent, with deep intronic variants behaving differently from splice-region ones. And a benchmark computed against transcript-level evidence answers a different question than one computed against clinically curated pathogenicity — the first asks whether splicing changed, the second whether the change mattered.
The reference set problem is familiar from the benchmarking piece in this series: a metric characterizes a model and an evaluation set jointly, and quoting it without the set attached makes it unevaluable. The splice-prediction case is the same argument applied to a score rather than an F1.
Low scores are not the safe direction
There is an asymmetry worth stating explicitly, because pipelines tend to treat a low score as reassurance.
The deep intronic sensitivity figures make the point quantitatively: at a 0.5 threshold, most spliceogenic variants beyond 50 bp from an exon were missed.8 A deep intronic variant scoring 0.3 is not evidence of no splicing effect; it sits in a region where the tool's sensitivity at conventional thresholds is poor.
The ClinGen recommendations encode this asymmetry structurally. BP7 — the code for a silent change with no predicted splicing impact — is explicitly not to be applied within donor/acceptor splice regions, because those regions carry a higher prior probability of harboring spliceogenic variants.4 A low score in a high-prior context does not license a benign call. The prior does work the score cannot do.
Pushing thresholds down to recover sensitivity has its own cost, and one group quantified it plainly. Having used a 0.011 threshold for a bespoke application, the authors explicitly cautioned against generalizing it, noting that evidence from experimentally confirmed splice-neutral variants indicates such a threshold yields a false-positive rate above 50%.10
What the score does not tell you, even when it is right
Suppose the prediction is correct and splicing genuinely is disrupted. Several questions that determine clinical significance remain open.
Which aberration occurs
Exon skipping, pseudoexon inclusion, partial intron retention, whole intron retention, and partial exon deletion are different events with different protein consequences. The maximum delta score does not distinguish them — which is precisely why extension tools exist. The SpliceAI-10k calculator was built to predict aberration type, inserted or deleted sequence size, reading-frame effect, and altered amino acid sequence, using all four delta scores and their positions across a 10 kb window rather than the single maximum.11 It reports ≥84% accuracy for pseudoexon and partial intron retention prediction.11 The information needed for typing is in the full output; it is discarded when the pipeline keeps only the maximum.
Whether the frame survives, and whether NMD applies
An in-frame exon skip that removes a non-essential domain and a frameshift that triggers nonsense-mediated decay are both "splice-altering." The authors of the 10k calculator make the point that reading-frame and translation effects are critical to predicting pathogenicity and to designing validation experiments.11 A score cannot supply them.
What fraction of transcript is affected
This is the sharpest gap, and it has been measured. In the validation of the 10k calculator, aberration-type predictions generally aligned with assay results, but the maximum SpliceAI score did not accurately predict the level of aberrant expression.12 A variant producing 10% aberrant transcript and one producing 100% may score similarly. For a recessive condition, for a gene with haploinsufficiency as the mechanism, or for a modifier allele, that distinction is the clinical question.
Whether the tissue in question splices that way
Predictions are made on sequence alone and are therefore tissue-agnostic, while splicing is markedly tissue-specific. A prediction that holds in one tissue may not describe the transcript pool in the tissue driving the phenotype.
| Question | Answered by the score? | What answers it |
|---|---|---|
| Is splicing likely altered? | Yes — this is what it estimates1 | The score, read against a threshold calibrated for the variant class and gene |
| Which aberration type? | No | Full four-score output with positions; extension tools such as SAI-10k-calc11 |
| Does the reading frame survive? | No | Predicted transcript structure and size; downstream translation analysis11 |
| What fraction of transcripts? | No — score does not track aberration level12 | RNA-seq from patient material; quantitative splicing assay |
| Is the mechanism consistent with the gene's disease model? | No | Gene-level curation; PVS1 decision tree reasoning |
| Does it happen in the relevant tissue? | No | Tissue-appropriate RNA studies |
Thresholds and performance figures cited here are specific to the studies, genes, and variant classes described in the text.
The counterpoint, which is substantial
None of this makes splice prediction unreliable. The evidence that it works is strong and should be stated as plainly as the caveats.
In the NF1 evaluation, SpliceAI achieved an AUC of 0.975 with sensitivity and specificity around 94% at its optimal cut-off, and significantly outperformed the MaxEntScan and Splice Site Finder combination on both AUC and specificity.6 Among 30 confirmed splicing variants in non-canonical intronic regions in that cohort, 100% were correctly predicted.6 Against curated minigene data, a threshold of 0.285 delivered 90% sensitivity and 90% specificity, and agreement on aberration type reached 87% when the full output was processed through SAI-10k-calc.7 The 10k calculator itself reports 95% sensitivity and 96% specificity for identifying splice-impacting variants across 1,212 SNVs with curated assay results.11
These are strong numbers, and they support exactly the use the tool was designed for: prioritization. The failure mode is not the model. It is the step where a prioritization score is treated as a classification, and where a threshold calibrated for one gene, variant class, or evidence direction is applied to another.
It is also worth crediting how carefully the field has handled this. The developers published three operating points rather than one. ClinGen calibrated the score to evidence strengths with likelihood ratios and specified where codes may not be applied. The gene-specific panels rederived their own thresholds. The vocabulary and the guardrails exist.
What this means for a pipeline
Keep the whole output, not the maximum
The four delta scores and their delta positions carry the aberration-type and size information that the maximum discards.11 A pipeline that annotates only max_DS has thrown away the basis for typing the event before a curator ever sees it. Retain all four scores and positions in the variant record.
Use the threshold that matches the question
State which threshold is applied and why: PP3 or BP4 direction, gene-specific VCEP specification where one exists, and the variant class. Where a ClinGen expert panel has published a gene-specific calibration, that supersedes the general recommendation for that gene.5 Do not carry one laboratory's 0.5 into another context because it appeared in a methods section.
Treat the indeterminate band as a result
Scores between the BP4 and PP3 thresholds support neither code.3,4 That is a legitimate output and should be reported as such rather than being resolved by a default. Silent binarization at 0.5 converts an honest "insufficient computational evidence" into an implicit benign lean.
Do not apply benign codes in high-prior regions
BP7 should not be applied in donor/acceptor splice regions.4 More generally, a low score in a context with elevated prior probability — canonical splice region, a gene with a known splicing disease mechanism, a deep intronic position where sensitivity is documented to be poor — is weak evidence, not reassurance.
Escalate to RNA where the score is load-bearing
When a splice prediction is the pivotal evidence for a classification, the prediction should be confirmed rather than trusted. RNA studies can confirm or refute predicted splice effects,2 and only RNA answers the fraction-of-transcript question the score is documented not to track.12 This is the concrete operational link to the WGS plus RNA-seq pairing: the prediction identifies the candidate, and the transcript data supplies the evidence.
Record the tool version and annotation set
Scores depend on the model version, on masked versus raw output, and on the gene annotation used — masked files are recommended for variant interpretation and raw files for alternative splicing analysis.1 A stored score without those parameters is not reproducible, which matters for reanalysis, as covered in the variant classification drift piece.
The shape of the problem
A model produces a continuous, well-characterized output. The output is compressed into a threshold call. The threshold was calibrated on a particular set of variants in a particular set of genes for a particular purpose, and none of that context is attached to the number when it moves downstream. A curator sees 0.9 and reads certainty; the model reported a probability that the splicing pattern differs from reference, which is a narrower claim than the one being made from it.
The recurring pattern in this series is a pipeline performing a resolution that is invisible in its output. Here the resolution happens twice: once when four scores and their positions collapse into a maximum, and again when a continuous probability collapses into a binary at a threshold nobody records. Both discard exactly the information a downstream interpretation needs, and neither leaves a trace in the variant record.
The tools are good. The published guidance is careful and specific. What travels onto the report is a number and a gene name.
References
- Illumina. SpliceAI documentation and delta score definition. GitHub / PyPI package documentation. https://github.com/Illumina/SpliceAI (developer documentation, cited for the score definition and the 0.2 / 0.5 / 0.8 characterization)
- SpliceAI precomputed splice impact scores — documentation on scope and limitations. https://www.helena.bio/docs/databases/spliceai-precomputed (third-party tooling documentation, cited for the functional-consequence limitation)
- Walker LC, Hoya M de la, Wiggins GAR, et al. Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicing: recommendations from the ClinGen SVI Splicing Subgroup. American Journal of Human Genetics. 2023;110(7):1046–1067. https://www.sciencedirect.com/science/article/pii/S0002929723002033
- Application of the ACMG/AMP framework to capture evidence relevant to predicted and observed impact on splicing: recommendations from the ClinGen SVI Splicing Subgroup (PMC deposit of ref. 3). https://pmc.ncbi.nlm.nih.gov/articles/PMC9980257/
- ClinGen Myeloid Malignancy Variant Curation Expert Panel. Specifications to the ACMG/AMP variant interpretation guidelines, version 2 (RUNX1). https://clinicalgenome.org/site/assets/files/7086/clingen_myelomalig_acmg_specifications_v2.pdf
- Performance evaluation of SpliceAI for the prediction of splicing of NF1 variants. Genes. 2021;12(9):1308. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8472818/
- Calibration of splicing predictions against minigene and patient-derived RNA results (SAI-10k-calc aberration-type agreement). https://www.researchgate.net/publication/330464045_Predicting_Splicing_from_Primary_Sequence_with_Deep_Learning
- Canson DM, Davidson AL, de la Hoya M, et al. SpliceAI-10k calculator for the prediction of pseudoexonization, intron retention, and exon deletion. Bioinformatics. 2023;39(4):btad179. https://academic.oup.com/bioinformatics/article/39/4/btad179/7109800 (cited for the deep-intronic sensitivity figures at 0.5 and 0.05 thresholds)
- Riepe TV, Khan M, Roosing S, Cremers FPM, 't Hoen PAC. Benchmarking deep learning splice prediction tools using functional splice assays. Human Mutation. 2021;42(7):799–810. https://pmc.ncbi.nlm.nih.gov/articles/PMC8360004/
- Dawes R, Bournazos AM, Bryen SJ, et al. SpliceVault predicts the precise nature of variant-associated mis-splicing. Nature Genetics. 2023;55:324–332. https://www.nature.com/articles/s41588-022-01293-8
- SpliceAI-10k calculator: performance and scope (PMC deposit of ref. 8). https://pmc.ncbi.nlm.nih.gov/articles/PMC10125908/
- Application of SAI-10k-calc to curated splicing assay data, including the observation that maximum SpliceAI score did not accurately predict level of aberrant expression. https://www.researchgate.net/publication/369856685_SpliceAI-10k_calculator_for_the_prediction_of_pseudoexonization_intron_retention_and_exon_deletion

