Multiple Testing Across Omics: Why the Genome-Wide Correction You Trust Breaks When You Integrate Data Types

Multiple Testing Across Omics: Why the Genome-Wide Correction You Trust Breaks When You Integrate Data Types
THE CORRECTION threshold = α / m m = number of tests GWAS assumes m ≈ 1,000,000 → 5 × 10⁻⁸ (effective, not raw, via LD) multi-omic breaks it: ▪ features correlate across layers ▪ same signal counted twice ▪ m is ambiguous ▪ tests aren't independent Too strict → miss real hits. Too loose → false ones. The count decides both. ZETOBIT · INSIGHT SERIES Multiple Testing Across Omics The genome-wide correction you trust breaks when you integrate data types. Kanna Nandakumar, PhD zetobit.com

Zetobit · Bioinformatics Insight Series

Multiple Testing Across Omics: Why the Genome-Wide Correction You Trust Breaks When You Integrate Data Types

The famous GWAS threshold of 5×10⁻⁸ is not a law of nature. It is a Bonferroni correction resting on a specific count of independent tests — and when you combine genomics, transcriptomics, proteomics, and more, that count becomes ill-defined and the correction quietly stops meaning what you think it does.

Anyone who has run a genome-wide association study knows the number: 5×10⁻⁸, the genome-wide significance threshold. It feels like a fixed constant of the field. It is not. It is a Bonferroni correction — the standard 0.05 significance level divided by the number of independent tests — built on the estimate that scanning the common-variant human genome amounts to roughly one million effectively independent tests.1 Divide 0.05 by a million and you land near 5×10⁻⁸. The threshold is only as sound as that count of one million.

What makes that count defensible within a single GWAS is a subtle and important point: the millions of SNPs tested are not independent. Linkage disequilibrium (LD) — the correlation between nearby variants — means the effective number of independent tests is far smaller than the raw number of markers.2 The whole apparatus of genome-wide correction is built on carefully estimating that effective number, accounting for the genome's correlation structure, rather than naively counting SNPs.1 This works because the dependence is understood and confined to one data type. Multi-omic integration removes both of those comforts.

Correction is a division problem, and the denominator is the trap

Every multiple-testing correction, whether it controls the family-wise error rate (Bonferroni, Holm) or the false discovery rate (Benjamini-Hochberg), turns on the number of tests m.3 Bonferroni multiplies each p-value by m (equivalently, divides the threshold by m); Benjamini-Hochberg ranks p-values against a sequence of m-dependent thresholds.4 Get m right and the error control is honest. Get it wrong and everything downstream is miscalibrated — too large a denominator makes the test needlessly conservative and buries real hits; too small makes it anti-conservative and lets false ones through.

In a single, well-understood assay, m is knowable. In multi-omics, the very question "how many tests am I doing?" loses a clean answer. You are testing SNPs, transcripts, proteins, metabolites, methylation sites — data types with different feature counts, different correlation structures, and, critically, correlations between them. The denominator is no longer a fact you can look up; it is a modeling choice, and different reasonable choices give materially different thresholds.

ONE BIOLOGICAL SIGNAL, COUNTED ONCE PER OMIC LAYER causal locus SNP (eQTL) transcript protein (pQTL) 3 "independent" tests — but 1 underlying effect Treating layers as independent inflates m (over-correct) OR ignores sharing (under-correct). On average ~3.5 molecular phenotypes trace back to a single GWAS locus.
A single causal locus can drive a correlated SNP, transcript, and protein signal. Counting each as a separate independent test misrepresents the true number of distinct hypotheses — inflating the denominator if you over-count, or, if you ignore the cross-layer sharing, letting one effect masquerade as independent corroboration.

The independence assumption fails in a new direction

Within one omic layer, dependence is already the subject of careful methodology — the LD-aware effective-test-count machinery exists precisely because SNPs correlate.2 Multi-omics adds a second, orthogonal kind of dependence: features in different layers that reflect the same underlying biology. A causal variant creates an eQTL (expression signal), which creates a pQTL (protein signal); a joint analysis of GWAS and molecular-QTL data across 50 traits found roughly 3.5 trait-associated molecular phenotypes per GWAS locus on average — vivid evidence that one genetic signal echoes across omic layers rather than generating independent ones.5

This breaks standard FDR procedures in both directions. Benjamini-Hochberg assumes independence (or a specific benign form of positive dependence); observations from large-scale testing are frequently dependent in ways that can degrade both estimation and the stability of results.6 When correlated cross-omic features are pooled and corrected as if independent, two failure modes appear. If you treat every layer's features as separate tests, you inflate m and over-correct, discarding real associations. If you treat a shared signal appearing in three layers as three independent confirmations, you fool yourself into false confidence — the corroboration is an artifact of measuring one effect three ways.

Correlated evidence is not independent evidence. The intuitive appeal of multi-omics is that a hit confirmed across genome, transcriptome, and proteome feels more trustworthy. Sometimes it is. But when those layers are causally linked, the "independent confirmation" can be a single biological fact reflected in three mirrors. Multiple-testing correction that ignores the linkage counts the mirrors, not the object.

Why "just Bonferroni everything" and "just BH everything" both fail

Pooling and Bonferroni-correcting across all omic features is simple and valid under arbitrary dependence — Bonferroni's guarantee doesn't require independence.3 But it pays for that robustness with severe conservatism: lumping hundreds of thousands of correlated features into one enormous m sets a threshold so strict that all but the largest effects vanish. You control the error rate by finding almost nothing.

Applying Benjamini-Hochberg across the pooled features recovers power, but its error guarantee rests on a dependence assumption that cross-omic correlation can violate, and it implicitly assumes the tests are exchangeable and similarly powered — untrue when a well-powered proteomic assay and a sparse metabolomic one are mixed into one ranking.7 Layers differ in sample size, signal-to-noise, and feature count, so a single flat correction systematically favors the higher-powered layer and starves the others.7

Neither default is wrong so much as mismatched to the structure of the data. The problem is not choosing between Bonferroni and BH; it is that both, applied flatly, pretend the multi-omic feature space is something it isn't.

What actually works

Handling multiple testing across omics honestly means respecting the structure rather than flattening it:

  • Use the effective number of tests, not the raw count. Within each layer, estimate independent tests from the correlation structure (as GWAS does with LD) rather than counting features; this alone corrects much of the over-conservatism.2
  • Correct hierarchically. Control error within each omic layer first, then across layers, so a low-powered layer isn't drowned by a high-powered one and the structure is respected at each level.6
  • Model cross-layer dependence explicitly. Use methods designed for integrative or feature-set testing that build the multiple-testing adjustment into the model and account for correlated features across omics, rather than bolting a flat correction on afterward.8
  • Distinguish shared from independent signal. Before treating a cross-omic hit as corroboration, test whether the layers share a causal signal (e.g., colocalization) versus contributing independent evidence — they demand different interpretations.5
  • Report the denominator and its basis. State how many tests were counted, how the effective number was derived, and how cross-layer dependence was handled, so a reader can judge whether the threshold means what it claims.3

As with the other blind spots in this series, the trustworthy analysis makes its assumptions legible: what m was, why, how dependence within and across layers was modeled, and whether shared signals were double-counted. A p-value threshold reported as if it were a universal constant — 5×10⁻⁸ carried unchanged into a multi-omic setting it was never derived for — looks rigorous but may be controlling nothing in particular.

The takeaway

Genome-wide significance is a triumph of careful multiple-testing correction — but a triumph built for one data type, resting on an LD-derived count of independent tests that does not transfer to integrated omics. Combine data types and the denominator becomes ambiguous, the independence assumption fails in a new cross-layer direction, and the same biological signal can be counted once, thrice, or not at all depending on choices most analyses never state. Correlated evidence across omics is not the same as independent confirmation, and a correction that can't tell the difference isn't protecting you. The number of tests was never just a number — and in multi-omics, pretending it is is where the errors hide.

References

  1. Revisiting the genome-wide significance threshold for common-variant GWAS (5×10⁻⁸ as Bonferroni for the effective number of independent tests under LD; ~1M independent tests). G3. 2021. academic.oup.com
  2. Calibrating genome-wide significance — effective number of independent tests is substantially lower than the total number of variants due to LD correlation structure. Sci Rep. 2025. nature.com
  3. Multiple testing correction in proteomics (FWER vs FDR; Bonferroni valid under arbitrary dependence but conservative; choice depends on number of tests and dependence structure). MetwareBio. 2026. metwarebio.com
  4. Visualizing the costs and benefits of correcting p-values for multiple testing in omics (Bonferroni multiplies by m; BH ranks against m-dependent thresholds; BH more sensitive). bioRxiv. 2021. biorxiv.org
  5. Joint analysis of GWAS and multi-omics QTL summary statistics — ~3.5 trait-associated molecular phenotypes per GWAS locus; signals shared across omic layers. Cell Genomics. 2023. cell.com
  6. Multiple testing approaches for hypotheses in integrative genomics — conventional FDR assumes independence/exchangeability; large-scale tests are often dependent, degrading accuracy and adding variability. WIREs Comput Stat. wiley.com
  7. MultipleTesting.com — BH/Storey q assume all tests have equal power; tests differ in sample size, signal-to-noise, and biology, motivating weighted/data-driven correction. PLOS One. 2021. journals.plos.org
  8. Multiple testing of mix-and-match feature sets in multi-omics — closed-testing/TDP framework with built-in multiple-testing correction across correlated cross-omic feature sets. PMC. PMC12825407
Previous
Previous

Phasing and Haplotype Assembly Limits: Why "Phased" Has a Resolution You Rarely See Reported

Next
Next

Batch Integration in Single-Cell RNA-Seq: Why Removing Technical Variation Can Erase the Biology You Came to Find