Phasing and Haplotype Assembly Limits: Why "Phased" Has a Resolution You Rarely See Reported

Phasing and Haplotype Assembly Limits: Why "Phased" Has a Resolution You Rarely See Reported
TWO HAPLOTYPES block 1 unphased gap block 2 — switch error everything downstream flips flip error (single site) N50 tells you block LENGTH. Not whether it's CORRECT. "phased" = broken into blocks, each with an error rate, separated by gaps. Rarely all reported together. ZETOBIT · INSIGHT SERIES Phasing & Haplotype Assembly Limits "Phased" has a resolution — switch errors, block boundaries — you rarely see reported. Kanna Nandakumar, PhD zetobit.com

Zetobit · Bioinformatics Insight Series

Phasing and Haplotype Assembly Limits: Why "Phased" Has a Resolution You Rarely See Reported

A phased genome sounds like a solved genome — two clean parental haplotypes, allele by allele. In reality "phased" means the chromosome was broken into blocks, each internally consistent but disconnected from its neighbors, each carrying an error rate, and each capable of flipping the entire downstream haplotype at a single wrong junction.

Phasing is the step that turns a list of heterozygous genotypes into haplotypes — assigning each allele to the maternal or paternal copy of the chromosome.1 It matters enormously: compound heterozygous variants (two hits in the same gene) are only interpretable if you know whether they sit on the same copy or opposite copies, and haplotype structure underpins imputation, eQTL work, and population genetics. When a report says a genome is "phased," it invites the mental image of two continuous, correctly separated chromosomes. That image is almost never what the data supports.

What "phased" actually delivers is a set of haplotype blocks: contiguous stretches the method declares internally phased, separated by gaps it could not resolve, each block carrying its own error rate — and, crucially, with no reliable phase relationship between blocks.2 The single headline number usually reported, block N50, describes only how long those blocks are. It says nothing about whether they are right.

Why the chromosome breaks into blocks

Read-based phasing works by finding reads (or read pairs) that span two or more heterozygous sites at once, physically linking their alleles. Where reads bridge consecutive variants, phase propagates; where they don't — across a stretch of low heterozygosity, a repetitive region, or simply a gap longer than the reads — the linkage breaks and a new block begins.3 Because this depends on read length reaching from one heterozygous site to the next, block structure is fundamentally limited by the sequencing technology and by how densely variants occur.

That second dependence has a consequence rarely stated: phasing contiguity varies with the genome being phased. In samples with higher heterozygosity, the denser spacing of heterozygous sites lets reads bridge more junctions, producing longer blocks; in samples with lower heterozygosity, blocks are shorter — a pattern visible as multi-fold differences in phase-block N50 between populations sequenced identically.4 "Phased to N50 of X" is therefore not a fixed property of a pipeline; it depends on the sample.

Switch errors: one wrong junction, everything after it inverted

Phasing errors are not all equal, and the distinction is the heart of the matter. A switch error is a point where the two haplotypes cross over: from that position onward, the entire remaining segment of the block is assigned to the wrong parental copy.5 A single switch doesn't corrupt one allele — it inverts every allele downstream of it within the block. A flip error is the more contained cousin: a single heterozygous site placed on the wrong haplotype against an otherwise correct background (equivalently, two switch errors in immediate succession that cancel out).5

The two have very different consequences. A flip mislabels one variant. A switch relocates a whole run of variants to the wrong chromosome copy — which is exactly the failure that breaks compound-heterozygote interpretation: two pathogenic variants that are truly in trans (opposite copies, disease-causing) can be made to look like they're in cis (same copy, often benign), or vice versa, by a single switch between them. In benchmarks, switch and flip errors occur at broadly comparable overall frequencies — roughly half and half of all errors, depending on method — so the more damaging switch type is not rare.6

SWITCH vs FLIP — SAME COUNT, DIFFERENT DAMAGE Truth A A A A A A A B B B B B B B Switch A A A B B B B B B B A A A A all 4 downstream sites now wrong Flip A A B A A A A B B A B B B B 1 site wrong, rest correct A switch flips cis/trans for every pair spanning it — the error that breaks compound-het calls.
Both are "one error," but a switch error reassigns everything downstream to the wrong haplotype, while a flip error misplaces a single site. For determining whether two variants are in cis or trans, a switch between them silently inverts the answer.

N50 measures length, not correctness — and they trade off

The most consequential reporting gap is that phasing is very often summarized by block length (N50) alone, yet N50, exactly as in genome assembly, does not represent the accuracy or quality of the result.2 You can make blocks longer by connecting variants more aggressively across weak evidence — and doing so introduces more switch errors. This is a genuine, measured tradeoff: relaxing the confidence threshold for joining variants increases block length by orders of magnitude while switch error climbs with it.7 A long N50 can therefore signal either excellent data or an aggressive setting that stitched blocks together through junctions it shouldn't have.

This is why length and accuracy must be reported together, and why quality-aware metrics exist: quality-adjusted block length (QAN50) breaks blocks at their error sites before computing contiguity, and the honest way to describe phasing is to state block N50 and switch error rate and flip error rate and the fraction of heterozygous sites actually phased.8 A single number cannot capture a result that lives in that four-dimensional space.

"Chromosome-scale phasing" often isn't read-based at all. Reaching block lengths that span whole chromosomes usually requires bringing in extra information — Hi-C proximity data, parental genotypes (trios), or population reference panels — because read length alone can't bridge every low-heterozygosity gap.3 Each source has its own failure mode: Hi-C introduces its own switch errors, trios need parental samples, and statistical/population phasing is a different beast entirely.

Statistical phasing trades physical evidence for a population model

When reads can't link variants, population-scale statistical phasing infers haplotypes from patterns of linkage disequilibrium across thousands of samples rather than from physical read linkage.9 It scales to biobank-sized cohorts and resolves common variants well — but it struggles precisely where interpretation often matters most: rare variants and singletons, which carry little population signal because they appear in too few people to triangulate.9 Modern methods have pushed this frontier impressively — one reports switch error rates below 5% even for variants seen in just one sample in 100,000 — but that residual error concentrates on the rare variants that clinical and functional work most often cares about, and the phase of a statistically inferred allele is a probabilistic claim, not a physically observed one.10

What a defensible phasing report requires

Treating "phased" as a resolution rather than a binary means reporting and interrogating the things a single N50 hides:

  • Report the full quality profile. Block N50 alongside switch error rate, flip error rate, and the percentage of heterozygous variants phased — never length alone.8
  • State the method and its evidence type. Read-based, Hi-C, trio, or statistical phasing fail differently; which one produced a given block determines how much to trust it.3
  • Check phase at the variants you care about. For a clinical compound-het call, confirm the two variants sit in the same block and assess the switch-error risk of the junction between them, rather than trusting a genome-wide average.6
  • Flag statistically phased alleles. Distinguish physically observed phase from population-inferred phase, especially for rare variants where statistical confidence is weakest.10
  • Don't chase N50 blindly. Recognize that longer blocks can mean more aggressive joining and more switch errors; the right operating point depends on whether you need contiguity or accuracy.7

As with the other blind spots in this series, the trustworthy result states its own resolution: how the genome was phased, into how many blocks of what length, with what switch and flip error rates, and how much was left unphased. A genome labeled simply "phased," with a single N50 and no error rate, looks complete but hides exactly the information needed to know whether any particular haplotype claim is real.

The takeaway

Phasing is indispensable and, at its best, remarkably good — but "phased" is not a yes/no property of a genome. It is a resolution, made of blocks with lengths, gaps between them, and switch and flip error rates within them, and the single number usually reported captures only one of those dimensions. A switch error can silently invert a compound-heterozygote call; a long N50 can hide aggressive over-joining; a statistically phased rare variant is a probability, not an observation. The discipline is to ask, of any phased result, not merely is it phased but at what resolution, by what evidence, and with what error — because the parts that don't get reported are exactly the parts that decide whether the haplotype is true.

References

  1. HapBridge — phasing assigns alleles at heterozygous sites to maternal/paternal haplotypes; read-length constraints produce localized phased blocks with unphased gaps. bioRxiv. 2025. biorxiv.org
  2. PhaseME — N50 of phase-block length does not represent accuracy/quality (as in assembly); phasing errors are flip (single SNV) and switch (incorrectly joined haplotypes). GigaScience / PMC. PMC7379178
  3. HapBridge — read-based phasing produces discrete blocks; low-heterozygosity/sparse-variant regions break phasing; Hi-C, parental genotypes, RNA-seq used to reach chromosome-scale. bioRxiv. 2025. biorxiv.org
  4. LongHap — phase-block N50 and switch error rate vary with heterozygosity: African samples (higher het) mean N50 ~1,727 kb vs East Asian ~321 kb, sequenced comparably. bioRxiv. 2026. biorxiv.org
  5. Improving population-scale statistical phasing — switch errors mis-phase entire contiguous segments; flip errors flip a single heterozygous genotype on a correct background. bioRxiv. 2023. biorxiv.org
  6. A Benchmark of Modern Statistical Phasing Methods — switch errors ~54% of all errors on average (47–61% by method); flips enriched at rare variants and CpG sites. bioRxiv. 2025. biorxiv.org
  7. Haplotype phasing in single-cell DNA-seq (HapCUT2) — lowering the phasing-score threshold increases block N50 by orders of magnitude with corresponding increases in switch error; explicit length/accuracy tradeoff. Bioinformatics. 2018. academic.oup.com
  8. Comparison of phasing strategies for whole human genomes — Quality-Adjusted N50 (QAN50) breaks blocks at each error site; reports SER alongside block length; a block is "declared phased" but may contain switches. PLoS Genet. 2018. journals.plos.org
  9. SAPPHIRE / population-scale statistical phasing — statistical methods use LD across many samples, excel at common variants, struggle at rare variants/singletons for lack of statistical information. bioRxiv. 2023. biorxiv.org
  10. Hofmeister RJ, et al. SHAPEIT5 — accurate rare-variant phasing; switch error rate below 5% for variants present in 1 of 100,000 samples; singletons less precise. Nat Genet. 2023. nature.com
Previous
Previous

Coverage Uniformity vs. Mean Depth: Why the Average That Reassures You Hides the Dropout That Fails the Sample

Next
Next

Multiple Testing Across Omics: Why the Genome-Wide Correction You Trust Breaks When You Integrate Data Types