Reading a Pathway Enrichment Figure: Why Twenty Bars Are Not Twenty Findings
Reading a Pathway Enrichment Figure: Why Twenty Bars Are Not Twenty Findings
The chart is a statement about the overlap between two lists, measured against a comparison set that was chosen for you and is almost never shown.
A slide arrives with a bar chart on it. Down the left, a column of biological processes with reassuringly meaningful names — immune response, cell cycle, oxidative phosphorylation. Along the bottom, a scale. The bars shorten as they descend, and the eye reads the chart the way it reads a leaderboard: these are the processes involved, in order of how involved they are.
That is not what the chart says. Each bar reports one thing only: that the gene list submitted to the software overlapped a stored list of genes more than would be expected by chance, given a comparison set of genes the analyst selected. It is a statement about two lists. Nothing in the calculation observed a protein, an activity, or a cell. Whether the process named on the axis is actually running differently in your samples is an inference laid on top, and it is usually the analyst's or the reader's rather than the software's.
Some of what follows is an unavoidable design property of the method. Some of it is a documented and very common mistake — this is an area where the published literature has been surveyed and the findings are not flattering. From where you sit, the distinction changes very little, because the questions you would ask are the same either way.
The bars are not independent of each other
Gene sets overlap heavily. Ontologies are hierarchical, so a term and its parent share members by construction; curated pathway databases carve overlapping territory out of the same well-studied biology. The consequence is that a handful of genes can light up a dozen terms at once. The current best-practice guidance is explicit about this and recommends that users check the overlap between their enriched sets — with an enrichment map or a clustered heat map — precisely because the observed enrichments may be driven by the same subset of genes.12
So the first correction to make when you look at one of these figures is arithmetic. Twenty bars are not twenty pieces of evidence. They may be three, or one. The way to find out is not to squint at the term names, which sound distinct by design, but to ask for the gene members behind each term and see how much they repeat. Every enrichment tool can emit that list, and it takes about a minute to read.
The comparison set you were not shown
An over-representation test asks whether your list is unusual relative to some universe of genes. That universe — the background, or reference list — is a parameter, and it changes the answer more than almost anything else on the slide.
The correct background for a transcriptomics experiment is the set of genes that could plausibly have appeared in the result at all: the ones detected in that tissue. Of roughly 78,000 genes annotated in the current Ensembl release, typically only 12,000–20,000 are expressed at detectable levels in any one tissue.1 Use the whole annotation instead, and every gene that was never expressed in that tissue is quietly counted as a gene that failed to appear in your list — which makes almost any tissue-specific list look enriched for the biology of that tissue.
The size of the effect is not subtle. In a worked example accompanying the 2026 best-practice guidance, an analysis run against the genes actually detected returned 207 significant pathways; the same analysis run against all annotated genes returned those plus a further 601 false positives.1 Three-quarters of the bars, in that instance, were an artifact of the comparison set.
That survey is the reason this question is worth asking out loud rather than assuming. The same screen found that 43% of analyses had not corrected p-values for multiple testing at all3 — in a procedure that runs thousands of tests at once. If a figure reports raw p-values, the number of bars on it is close to meaningless.
What the length of the bar measures
Almost every enrichment figure plots significance on its x-axis, which means the ordering rewards large, generic gene sets. A term containing four thousand genes needs only a mild excess to clear an adjusted threshold; a term containing twenty-four needs a dramatic one. Prioritizing results by p-value alone therefore emphasizes broad categories with moderate enrichment and pushes the small, specific, strongly enriched sets — usually the more interesting leads — off the bottom of the chart.1 Term size is itself known to skew these statistics independently of the underlying biology.4
There is a second-order problem here, which is that some widely used tools do not report an effect size at all.1 If the figure came from one of those, no one involved has seen a fold enrichment for those terms — it is not that the analyst chose not to plot it. Asking for it may mean asking for the analysis to be exported differently, which is a small cost and worth it.
Readers of the previous piece in this series will recognize the shape of this: the same conflation of confidence with magnitude that governs a differential expression table governs the figure built on top of it. The enrichment chart also inherits every property of that table, including which genes were tested and which were silently dropped, since the gene list it consumed came from there.
The database has a version, and a bias
The gene sets themselves are a curated body of literature, not a measurement, and they are unevenly complete. In an analysis of a decade of Gene Ontology releases, 58% of annotations belonged to just 16% of human genes, and enrichment results computed from early and later ontology versions on the same 104 gene signatures showed low consistency.5 Two things follow. The figure has a date: rerun in two years, against the same data, it will look different. And the terms that appear are drawn preferentially from biology that has already been studied, which is why a novel finding so often lands in a category that sounds slightly beside the point. Absence of a process from the chart is frequently a statement about the annotation, not about the sample.
None of this makes the figure worthless. It is a genuinely efficient way to compress a list of several hundred genes into a handful of themes worth pursuing, and the best-practice literature is clear that this is what it is for: generating hypotheses that then require separate validation.1 The failure mode is not the chart. It is the sentence that gets written underneath it — that a pathway was activated, that a mechanism was identified — when what was established is that one list overlapped another list more than a chosen comparison set would predict.
Three questions to ask
- What background list was used — the genes detected in this experiment, or the whole annotation? This single choice can multiply the number of significant terms severalfold. If the answer is "the default," the figure needs rerunning before it means anything.
- How much do the top terms overlap? Can I see the genes behind each one? Ask for the member genes or an enrichment map. Terms that share most of their genes are one result, and should be described as one result.
- What is the effect size for each term, and which database version was queried? Fold enrichment or enrichment score tells you what the bar length does not. The database version tells you how much of this is reproducible next year.
References
- Bora A, McKenzie M, Ziemann M. Ten common mistakes that could ruin your enrichment analysis. PLOS Computational Biology. 2026;22(4):e1014122. journals.plos.org
- Reimand J, Isserlin R, Voisin V, et al. Pathway enrichment analysis and visualization of omics data using g:Profiler, GSEA, Cytoscape and EnrichmentMap. Nature Protocols. 2019;14(2):482–517. nature.com
- Wijesooriya K, Jadaan SA, Perera KL, Kaur T, Ziemann M. Urgent need for consistent standards in functional enrichment analysis. PLOS Computational Biology. 2022;18(3):e1009935. journals.plos.org
- Karp PD, Midford PE, Caspi R, Khodursky A. Pathway size matters: the influence of pathway granularity on over-representation (enrichment analysis) statistics. BMC Genomics. 2021;22(1):191. bmcgenomics.biomedcentral.com
- Tomczak A, Mortensen JM, Winnenburg R, et al. Interpretation of biological experiments changes with evolution of the Gene Ontology and its annotations. Scientific Reports. 2018;8:5115. nature.com

