Polygenic Score Portability Across Ancestries: Why Accuracy Decays Continuously and a Population Label Cannot Tell You Where

Polygenic Score Portability Across Ancestries: Why Accuracy Decays Continuously and a Population Label Cannot Tell You Where — Zetobit
Zetobit Bioinformatics Insight Series
BIOINFORMATICS INSIGHT SERIES Polygenic Score Portability Why accuracy decays continuously, not in discrete groups training data centre accuracy genetic distance → Kanna Nandakumar, PhD · Zetobit

Statistical Genomics & Health Equity

Polygenic Score Portability Across Ancestries: Why Accuracy Decays Continuously and a Population Label Cannot Tell You Where

A polygenic score trained in one population predicts poorly in another — that much is well known. What is less appreciated is that the loss is not a step change between labeled groups. It is a smooth decline with genetic distance from the training data, and it happens inside a single ancestry label as surely as across the boundary between two.

Polygenic scores have become one of the most visible products of human genomics. They sum the small effects of thousands to millions of variants into a single number estimating a person’s genetic predisposition to a trait or disease, and they are increasingly proposed for the clinic — cardiovascular risk stratification, breast cancer models, statin targeting. The appeal is obvious: a genome you already sequenced, converted into an actionable risk estimate at almost no marginal cost.

The problem that shadows every one of these applications is portability. A score’s predictive accuracy is a property of the population it was trained in, and it does not travel. Because the great majority of genome-wide association data comes from participants of European ancestry — a well-documented and slow-moving imbalance — scores built on that data predict least well for everyone else, and the gap is large enough to matter clinically. But the usual framing of that gap — European versus non-European, one accuracy per group — hides the more important and more actionable structure underneath it.

The scale of the gap is not subtle

Start with the aggregate numbers, because they set the stakes. The Eurocentric bias in the underlying data is stark: individuals of European descent make up roughly 16% of the global population but account for approximately 79% of all GWAS participants.1 Scores built on that foundation lose accuracy as they move to less-represented groups, and the canonical estimate is that prediction accuracy (R²) can be on the order of 4.5-fold lower in individuals of African ancestry than in Europeans,1 a gap first quantified in the work that put polygenic-score health disparities on the map.3 A risk model that is four to five times less accurate in one group than another is not a tuning problem; deployed naively, it is a mechanism for widening the very health disparities precision medicine is meant to narrow.

The mechanical causes are understood. Effect sizes estimated in one population transfer imperfectly to another because of differences in allele frequency (a variant common enough to estimate well in the training group may be rare elsewhere), differences in linkage disequilibrium (the tag SNP on the array may be correlated with the true causal variant in one population but not another), and genuine effect-size heterogeneity from gene–gene and gene–environment interaction. Every one of these grows as the target population diverges from the training population.

The reframe: distance is continuous, labels are not

Here is where the standard picture misleads. Portability is almost always reported as one accuracy number per discrete ancestry cluster — European, African, East Asian, South Asian — as though each group were internally uniform. A large 2023 analysis using the UCLA ATLAS biobank (n = 36,778) and the UK Biobank (n = 487,409) showed that this is the wrong unit of analysis. Polygenic score accuracy decreases individual-to-individual along a continuum of genetic ancestry, in every population considered, even within traditionally labeled ‘homogeneous’ groups.2

The organizing quantity is genetic distance — essentially how far an individual sits from the centre of the training data when both are projected into the same principal-component space. Averaged across 84 traits, the correlation between that genetic distance and polygenic-score accuracy was −0.95: accuracy falls almost linearly as an individual moves away from the training cloud.2 The decline does not wait for a population boundary. Applying scores trained on white British individuals to individuals of European ancestries in ATLAS, those in the farthest genetic-distance decile had 14% lower accuracy than those in the closest decile — a substantial spread within a single conventional label.2

The discrete view one accuracy per labeled group acc. EUR SAS AFR The continuous view accuracy falls with genetic distance acc. genetic distance → groups overlap — the label isn’t the unit
Same data, two units of analysis. Reporting one accuracy per ancestry cluster (left) treats each group as internally uniform and draws hard lines between them. Modelling accuracy against genetic distance (right) reveals a continuous decline whose within-group spread overlaps across labels — the closest individuals of one group can be better predicted than the farthest individuals of the group next door.

The overlap is the practical punchline. In the same analysis, the closest genetic-distance decile of individuals with Hispanic/Latino ancestries showed polygenic-score performance similar to the farthest decile of individuals with European ancestries.2 A label-based rule that grants full confidence to everyone in the European cluster and discounts everyone in the Hispanic/Latino cluster gets both of those subgroups wrong. The distance is doing the work the label was being asked to do — and doing it individual by individual.

A population label is a lossy compression of an individual’s position in genetic space. It draws a boundary where the underlying quantity — distance from the training data — varies smoothly across it. Two people with the same label can sit at opposite ends of the accuracy range, and two people with different labels can sit side by side.

Why the label is not just imprecise but unstable

There is a second, subtler problem with cluster-based reporting: the clusters themselves are not fixed objects. Assigning an individual to a discrete genetically inferred ancestry depends on the algorithm and the reference panel used, and different choices can place the same person in different clusters, which then yields a different reported accuracy for that person.2 Worse, a meaningful fraction of people are not confidently assignable to any cluster at all — in ATLAS, about 6% of individuals could not be placed into a genetic-ancestry group given the reference panels used, and under a label-only framework those people fall outside polygenic-score characterization entirely.2 A continuous genetic-distance measure has no such gap: everyone has a distance, including the admixed and the unassignable, so everyone can be given an honest, individualized accuracy.

The estimate itself can be biased, not just noisier

It would be reassuring if distance only widened the error bars — if a far-from-training score were merely less certain but still centred on the truth. The 2023 analysis found something more troubling. Across 84 traits, genetic distance was significantly correlated with the polygenic-score estimates themselves for 82 of them, and for 30 traits that correlation ran in the opposite direction to the correlation between distance and the actual measured trait.2 In other words, the score can drift systematically with ancestry in a way the real phenotype does not — a bias, not just added variance. The paper’s neutrophil-count example is instructive: a variant that strongly lowers neutrophil counts in individuals of African ancestry is absent from a European training set, so the score moves the wrong way relative to the true trait for exactly the group already least well served.2

What this means for anyone deploying a score

The takeaways are less about a single fix and more about how to report and interpret honestly:

  • Report accuracy against genetic distance, not just per cluster. A single R² per labeled group hides a spread wide enough to flip the ranking between adjacent groups. Distance gives every individual, including admixed and unassignable ones, a personalized reliability.2
  • Treat a European-derived score as calibrated only near its training centre. Accuracy is highest for individuals closest to the training data and declines smoothly outward — the ~4.5-fold aggregate gap for African-ancestry individuals is the endpoint of that decline, not a special case.1
  • Diversify the training data, because it helps everyone. Broadening training beyond European ancestries improves effect-size estimation for variants common outside Europe and, by pulling the training centre toward more people, shortens the genetic distance for target individuals across the board.2 Multi-ancestry methods such as PRS-CSx exist precisely to exploit this.
  • Watch for directional bias, not only lost precision. Because a score can drift opposite to the true trait far from training, a confidently-reported extreme score in a distant individual may be an artifact of missing population-specific variants rather than real risk.2

The deeper lesson is conceptual. Human genetic variation is a continuum, not a set of discrete bins, and a polygenic score inherits that geometry whether or not the report acknowledges it. The population label was always a convenient stand-in for the thing that actually governs accuracy — where you sit relative to the data the model learned from. Once you measure that directly, the discrete gap dissolves into a gradient, and the gradient is both more honest and more useful: it tells you not just that a score travels poorly, but exactly how far it has travelled for the person in front of you.

References

  1. Polygenic risk scores: navigating the future of precision medicine through economic, ethical, and scientific advancements. iScience. 2025;28. Summarizing Martin AR, et al. Nat. Genet. 2019;51:584–591. cell.com/iscience/fulltext/S2589-0042(25)02636-7
  2. Ding Y, Hou K, Xu Z, et al. Polygenic scoring accuracy varies across the genetic ancestry continuum. Nature. 2023;618(7966):774–781. nature.com/articles/s41586-023-06079-4
  3. Martin AR, Kanai M, Kamatani Y, et al. Clinical use of current polygenic risk scores may exacerbate health disparities. Nature Genetics. 2019;51:584–591. nature.com/articles/s41588-019-0379-x
© Zetobit LLC · Bioinformatics Insight Series zetobit.com
Previous
Previous

HLA Typing from Short Reads: Why the Most Polymorphic Region in the Genome Defeats Standard Alignment

Next
Next

Calling Mitochondrial Heteroplasmy: Why a Low-Level Variant Is as Likely to Be Nuclear DNA as Real