Peak-to-Gene Assignment: Why the Nearest Gene Is a Guess That Enters the Table as a Fact

Peak-to-Gene Assignment: Why the Nearest Gene Is a Guess That Enters the Table as a Fact
Zetobit · Bioinformatics Insight Series
ZETOBIT · BIOINFORMATICS INSIGHT SERIES Peak-to-Gene Assignment Why the nearest gene is a guess that enters the table as a fact BEST BENCHMARKED MODEL, K562 70% precision DISTANCE ALONE 32% Neither number appears in the gene column. Kanna Nandakumar, PhD zetobit.com
Regulatory Genomics

Peak-to-Gene Assignment: Why the Nearest Gene Is a Guess That Enters the Table as a Fact

Every accessibility result becomes a gene list at some point, and the step that converts one to the other is usually a default nobody chose. It carries no score, and everything downstream inherits it.

An ATAC-seq analysis produces coordinates. A report produces gene names. Somewhere between the two sits a join, and in most pipelines that join is proximity: assign the element to the closest transcription start site, or compute a per-gene score by summing nearby peaks with a distance weight. Both are defaults in the standard toolkits. Neither carries a confidence value, an alternative candidate, or any record that a choice was made. From that point on the analysis is about genes, and the only thing that was actually measured — where the fragments landed — stops being mentioned.

Three assignment rules are in common use: proximity, correlation between accessibility and expression across cells, and activity-and-contact models trained on perturbation data. This piece is about what each rule can and cannot establish. The separate statistical question of how many independent observations a correlation across metacells actually rests on is a different problem, and belongs in its own discussion.

First, a correction to the folklore

The usual justification for abandoning nearest-TSS is that enhancers routinely skip genes. That claim traces mostly to chromosome-conformation data, where it is dramatic: an early ENCODE analysis of promoter looping reported that only around 7% of interactions involved the nearest gene.9,10 Repeated often enough, this became the received wisdom that proximity is close to useless.

The perturbation data does not support that reading, and the distinction matters. The ENCODE enhancer-gene effort made a deliberate methodological choice to use only gold standards describing regulatory interactions — derived from genetic perturbations — rather than physical interactions measured by Hi-C or ChIA-PET.1 On that footing, the picture is much more local: across 352 biosamples, a median of 31.8% of predicted regulatory interactions occurred within 10 kb and 86.8% within 100 kb, consistent with the distance distributions of CRISPR perturbations, eQTLs and GWAS positive controls.1 A recent analysis argues directly that mammalian enhancers act proximally and seldom skip active genes.10

Contact maps and perturbation maps disagree about skipping because they are not measuring the same claim. Proximity predicts physical contact poorly and regulatory effect rather well — just nowhere near well enough to be used without a score.

So the case against nearest-TSS is not that the answer is usually far away. It is that proximity is strongly informative and decisively insufficient, and the assignment step reports it as though it were neither.

How well does any rule actually do?

The K562 benchmark gives hard numbers. A harmonised gold standard of 10,411 element-gene pairs tested by CRISPR perturbation — reanalysed and harmonised from several published screens3,5 — yielded 472 positives against 9,938 negatives with good power, so roughly one tested pair in twenty-two is a real regulatory link.1 Against that, the supervised ENCODE-rE2G model using cell-type-specific DNase alone achieved 54% precision at 70% recall. An extended model adding H3K27ac, cell-type-specific Hi-C and ChIA-PET reached 70%. The ABC model on DNase plus averaged Hi-C reached 46%.1,2 Using an inverse function of distance in place of measured contact dropped precision at the same recall to 31.8%. Simple baselines — assigning each element to the closest expressed gene, or correlating element and promoter accessibility across cell types — performed less well than the modelled predictors.1

Read those figures the right way round. The best available model, given a full complement of assays and supervised training on perturbations in the very cell type being predicted, is wrong about three predictions in ten at a useful recall. Nearest-TSS sits below the distance-only baseline, because it discards even the graded information distance provides, and it emits no score to discount.

One detail from the same benchmark deserves more attention than it gets from people running ATAC. Substituting assays into the ABC model, ATAC-seq performed worse than DNase-seq as the enhancer-activity input — 41% versus 52% precision at 70% recall.1 The assay most laboratories actually run is the weaker input, so published accuracy figures for linking methods are, for a typical ATAC project, optimistic before anything else goes wrong.

The assignment is a property of the cell type, not of the genome

This is the part that makes proximity annotation structurally wrong rather than merely inaccurate, and it is the strongest argument in the literature against a coordinates-only gene column.

In the ENCODE encyclopedia of 13.5 million regulatory interactions, an element regulated a median of 1 gene but a mean of 1.6, and a gene had a median of 3 regulatory elements.1 Interactions were highly context-specific: 26.4% were detected in only one of 352 biosamples. And among elements predicted to regulate more than one gene, 64% were linked to genes in mutually exclusive sets of biosamples — the same coordinate, a different target, in a different tissue. For fine-mapped GWAS variants falling in enhancers, the average number of linked genes was 3.4 with a maximum of 56, and 62% were linked to different genes in different cell types.1,3

A gene name computed from coordinates alone is therefore not a noisy estimate of a fixed quantity. It is an answer to a question that has no cell-type-independent answer, presented as though the mapping were a constant of the genome. That is a different kind of error, and no amount of care with the distance cutoff addresses it.

The errors are not spread evenly

Two features of the biology concentrate the failures where they do the most damage. First, promoters differ in how much they use distal elements at all: ubiquitously expressed genes were five-fold less likely to have distal regulatory elements in the K562 CRISPR data,1,8 which means distal enhancers are disproportionately the regulators of cell-type-specific genes. Those are precisely the genes a differential accessibility experiment is usually about. The rule performs worst on the biology under study.

Second, gene density decides the outcome. Where an element sits among several transcription start sites, the winner is chosen by base pairs, and multi-center CRISPRi analysis found that 86% of significant regulatory elements were within the same TAD as their target gene — a domain typically containing several genes, all plausible, only one correct.4 Because gene density is not independent of gene function, a pathway enrichment computed on a proximity-derived list inherits a structured bias rather than random noise.

Correlation as an assignment rule

Linking peaks to genes by correlating accessibility with expression across cells — the approach behind Signac's peak-gene linking and ArchR's peak-to-gene links6,7 — is a real improvement over coordinates, because it uses data that varies with the biology rather than a fixed genomic property. Judged purely as an assignment rule, it has two limits worth stating plainly.

It can only find a link where both the element and the gene vary across the cells profiled. An enhancer that is constitutively active in every cell in the experiment cannot be linked to anything, and its absence from the output is a detection limit rather than evidence that it regulates nothing. And co-variation separates a regulator from a bystander only when the bystander fails to co-vary — which, for genes sharing a regulatory domain and a cell-type program, is often not the case. In the ENCODE benchmark, cross-cell-type correlation of element and promoter accessibility was among the baselines that underperformed the modelled predictors.1 Both limits are about what the rule can establish, and both survive regardless of how the significance of the correlation is computed.

Assignment ruleEstablishesDoes not establishDefensible when
Nearest TSS or distance-weighted gene score That a gene is close by Regulation, target identity in gene-dense regions, or any cell-type context Used as an exploratory label and never reported as a finding
Accessibility–expression correlation Co-variation across the cells actually profiled Causal direction; anything about invariant elements, which cannot be linked at all The variation requirement and the distance window are both stated
Activity-and-contact models (ABC, ENCODE-rE2G) A calibrated probability benchmarked against perturbations The link in your cell type, unless the inputs came from your cell type The score is carried forward rather than thresholded away silently
CRISPRi or equivalent perturbation A regulatory effect on that gene in that cell type Generality to other cell types or conditions Always — this is the only direct evidence in the list

Every rule above the last one produces the same artefact: a gene name in a column. The rules differ in what that name means, and the column has no field to record which rule produced it.

ONE ELEMENT, THREE CANDIDATES Cell type A element GENE 1 GENE 2 GENE 3 nearest TSS → assigned perturbation → actual target SAME COORDINATE, DIFFERENT TISSUE Cell type B GENE 1 GENE 2 GENE 3 actual target here
Assigning an element to the closest transcription start site picks a candidate; a perturbation identifies a target. In the ENCODE maps, 64% of elements predicted to regulate more than one gene did so in mutually exclusive sets of cell types, so the correct answer changes with the tissue while the coordinates do not. A gene column computed from distance alone encodes the assumption that the mapping is fixed.

What to do about it

  1. Never let a gene name appear without the rule that produced it. “Nearest TSS within 100 kb” belongs in the figure caption, not in the methods appendix.
  2. Carry the score, not just the link. Modelled predictions come with a probability; thresholding it away and keeping the name discards the only quantitative thing in the step.
  3. Use an activity-and-contact model where the inputs allow it, and use inputs from your cell type — the models are calibrated on cell-type-specific data and degrade without it.
  4. Discount published accuracy for ATAC input. ATAC underperformed DNase as the activity measurement in the same benchmark; assume you are below the quoted figure.
  5. Keep more than one candidate. The median element regulates one gene, but the mean is higher and the tail is long; a single-name column throws away the ambiguity that actually exists.
  6. Treat a link established in another cell type as a hypothesis, not an annotation. Two-thirds of multi-gene elements switch targets between tissues.
  7. State the distance window explicitly. It is a hard cap on what can be found, and nothing outside it can ever be reported as absent.
  8. For correlation-based links, state the variation requirement and do not read a missing link as evidence that an element regulates nothing.
  9. Where a gene name carries real weight — a target nomination, a client deliverable, a go/no-go — plan the perturbation. It is the only rule in the table that establishes what the report claims.

The join is the least examined step in the analysis

Most entries in this series turn on a number that means less than it looks like, or a word that upgrades a claim. This one turns on an operation. Peak-to-gene assignment is, mechanically, a join between two tables, and joins are the least scrutinised step in any pipeline because they feel like bookkeeping rather than inference. Alignment has quality scores. Peak calling has a q-value. Differential testing has an adjusted p-value. The join has a column of gene symbols, and the column is the input to everything a reader will actually remember.

None of this is an argument against linking peaks to genes, which is the entire point of running the assay. Proximity is genuinely informative — more so than the contact-map folklore suggests. The field has built better rules, benchmarked them honestly against genetic perturbations rather than against physical proximity, and published scored maps across hundreds of cell types that cost nothing to use. The situation is unusually good compared with most problems this series covers. Defaulting to nearest-TSS in 2026 is a choice rather than a constraint, and the choice is defensible only when it is visible.

The peak was measured. The gene was assigned. Only one of those is data, and only one of them appears in the figure.

References

  1. Gschwind AR, Mualim KS, Karbalayghareh A, et al. An encyclopedia of enhancer-gene regulatory interactions in the human genome. bioRxiv. 2023:2023.11.09.563812. doi:10.1101/2023.11.09.563812 (Preprint at time of writing; check for the peer-reviewed version.)
  2. Fulco CP, Nasser J, Jones TR, et al. Activity-by-contact model of enhancer–promoter regulation from thousands of CRISPR perturbations. Nature Genetics. 2019;51(12):1664–1669. doi:10.1038/s41588-019-0538-0
  3. Nasser J, Bergman DT, Fulco CP, et al. Genome-wide enhancer maps link risk variants to disease genes. Nature. 2021;593:238–243. doi:10.1038/s41586-021-03446-x
  4. Yao D, Binan L, Bezney J, et al. Multicenter integrated analysis of noncoding CRISPRi screens. Nature Methods. 2024;21:723–734. doi:10.1038/s41592-024-02216-7
  5. Gasperini M, Hill AJ, McFaline-Figueroa JL, et al. A genome-wide framework for mapping gene regulation via cellular genetic screens. Cell. 2019;176(1–2):377–390.e19. doi:10.1016/j.cell.2018.11.029
  6. Stuart T, Srivastava A, Madad S, Lareau CA, Satija R. Single-cell chromatin state analysis with Signac. Nature Methods. 2021;18(11):1333–1341. doi:10.1038/s41592-021-01282-5
  7. Granja JM, Corces MR, Pierce SE, et al. ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis. Nature Genetics. 2021;53(3):403–411. doi:10.1038/s41588-021-00790-6
  8. Bergman DT, Jones TR, Liu V, et al. Compatibility rules of human enhancer and promoter sequences. Nature. 2022;607:176–184. doi:10.1038/s41586-022-04877-w
  9. Sanyal A, Lajoie BR, Jain G, Dekker J. The long-range interaction landscape of gene promoters. Nature. 2012;489:109–113. doi:10.1038/nature11279
  10. Mammalian enhancers and GWAS variants act proximally and seldom skip active genes. bioRxiv. 2024:2024.11.29.625864. doi:10.1101/2024.11.29.625864 (Preprint; cited for its survey of contact-based skipping estimates and its counter-argument.)
Zetobit, LLC · Bioinformatics consulting · Lexington, KY · zetobit.com
Related in this series: Sparsity in Single-Cell ATAC-seq, Cell Composition in Bulk ATAC-seq, Selection Bias in Enrichment Analysis.
Previous
Previous

QC Thresholds as Sample Selection: Why the Cells That Failed Are Not a Random Draw

Next
Next

Cell Composition in Bulk ATAC-seq: Why a Differentially Accessible Peak May Record Which Cells Were There