Sparsity in Single-Cell ATAC-seq: Why a Zero Is a Sampling Outcome, Not a Closed Locus
Sparsity in Single-Cell ATAC-seq: Why a Zero Is a Sampling Outcome, Not a Closed Locus
A scATAC-seq count matrix has the same shape as a scRNA-seq one and is handled by the same tools. Its zeros mean something different, and nearly every conclusion drawn from the matrix inherits that difference.
A single-cell ATAC-seq experiment produces a matrix. Rows are cells, columns are peaks, entries are counts. It loads into the same objects as a transcriptomic matrix, receives the same dimensionality reduction, and produces the same UMAP. The visual grammar is identical to single-cell RNA-seq, and so is the habitual reading: a zero means the feature was not detected, and a cell with a zero probably lacks the thing. In transcriptomics that reading is wrong in a bounded way. In chromatin accessibility it is wrong in an unbounded way, for a reason that sits in the molecule rather than in the pipeline.
This series has already covered ATAC-seq artifact regions — the blacklist argument, about places where signal appears that reflects the reference assembly rather than the chromatin. That is a false-presence problem, and it has a clean remedy: exclude the regions. What follows is the opposite failure. It is structural absence, it occurs at every region in every cell, and it cannot be excluded, because it is most of the matrix.
The arithmetic ceiling
An expressed gene is present in a cell as many transcript molecules, so a scRNA-seq count has room to vary. Chromatin accessibility is measured on DNA, which in a diploid cell exists in two copies at any autosomal locus. Whatever the sequencing depth, one cell can yield only a small number of transposition events at one site. The consequence is measured rather than theoretical: only about 1–10% of the peaks expected to be accessible are detected in any given cell, against 10–45% of expressed genes in scRNA-seq.2 Inverted, that is the sentence worth keeping: in a typical experiment, somewhere between 90% and 99% of the open chromatin in a cell produces nothing at all.
The matrix reflects it. Non-zero entries run at roughly 3% of a peak-by-cell matrix,6 and among the entries that exist, 90–95% are either 0 or 1. The mean of the non-zero counts in a cell rarely rises above 1.2 even in deeply sequenced cells — on average 62.8% lower than the equivalent figure in scRNA-seq.1 The data look almost binary. It is worth being precise about why: not because chromatin is binary, but because the sampling is thin enough that a second observation at the same site in the same cell is rare.
None of this is a protocol failure, and the budget arithmetic makes that clear. Open chromatin occupies roughly 1–2% of the genome, a target of 30–60 megabases, while 3′-biased scRNA-seq is aiming at something closer to 6 megabases. Matching the coverage would require scATAC-seq libraries about ten times larger than scRNA-seq libraries. In practice they are about twice as large: in one matched 10x multiome sample, a median of 10,000 fragments per cell against 5,000 UMIs.5 The assay is asked to cover ten times the territory on twice the budget. Sparsity is the designed operating point, not a defect in a particular run.
How much information is actually there
The question that matters is not how many zeros there are but whether the counts can distinguish a closed locus from an unsampled one. A recent analysis answered it in an unusually clean way.1 The authors built a hierarchical model of the generative process — a per-cell open or closed state, a Poisson number of true transposition events, and a binomial observation step for what survives to be sequenced — then simulated data from it and computed the posterior probability of the open state using the ground-truth parameters. That design is the important part. It is not a benchmark of one tool against another; it is an upper bound on what any method could achieve if it estimated every parameter perfectly.
The best cell in the parameter grid reached a mean AUROC of 0.84. Most of the grid was far worse. Matching real datasets against the simulations, fewer than 25% of peaks in most datasets carried enough counts to reach even AUROC 0.55 — meaning that more than 75% of features hold essentially no information about whether an individual cell is open at that site. A protocol engineered for higher Tn5 sensitivity raised that figure to 34%.1,10 The improvement is real and the direction is right; the ceiling is still low.
The authors draw a distinction the field should adopt: physical resolution versus informational resolution. The physical resolution is genuine — those fragments came from one cell, and nothing about the sparsity changes that. The informational resolution is a different question, and at a single locus in a single cell the answer is usually that there is nothing to conclude.
The zeros do not go away — they go into the normalization
Here is where the problem stops being a caveat and starts steering results. A standard workflow does not ignore the zeros; it normalizes them, and the normalization is largely made of them. The term-frequency half of TF-IDF divides each count by the cell's total, which is counts-per-million under a different name. When almost every entry is 0 or 1, the numerator barely varies and the denominator does the work — so after transformation, the largest source of variation between cells is their sequencing depth.1 Binarizing first, as several popular packages do by default, makes it worse by forcing every non-zero entry to the same value.
The effect is measurable. In a standard PBMC dataset, the first LSI component correlated with library size at −0.95.1 This is the origin of an instruction that appears in nearly every scATAC-seq tutorial: drop the first LSI component before clustering.9 The step is so routine that its justification has faded from view. It exists because the normalization did not work. Benchmarking confirms that TF-IDF is often ineffective at removing library-size effects,7 and using TF-IDF values for differential accessibility improves concordance with matched bulk data while simultaneously increasing false discoveries.8 A count-based model with an explicit observation probability, by comparison, produced a first component correlating with library size at only −0.34, with better silhouette widths and no component removal required.1
The routine fix is also not free, and this is the part that deserves more attention than it gets. The depth artifact and the biology are partly the same variable. A cell's total fragment count reflects how much of that cell was sequenced, which is technical, and also how many regions are open in that cell type at all, which is not. Deleting the component that tracks depth deletes some of each, and the output carries no record of the ratio.
Binarization hides an artifact rather than removing one
The other common response is to stop asking the counts to mean anything: if the ceiling is two, treat the entry as open or closed and move on. It sounds like the conservative choice. It was tested directly, and it is not.3 Binarization improved neither goodness of fit, nor clustering, nor cell-type identification, nor batch integration, and modeling fragment counts as Poisson-distributed outperformed it. The counts do carry information — nucleosome turnover happens on a timescale comparable to the tagmentation reaction, and the two alleles of a locus need not be in the same state.
One detail from that work should change how a matrix is read on receipt. Standard pipelines, including the vendor's, count transposase insertion ends rather than whole fragments. Because the reads are paired, this makes even counts predominate; odd counts arise mainly when one end of a fragment falls outside the peak.3 The distribution therefore has a shape imposed by a counting convention before any biology enters. Paired-insertion counting resolves it.4 Note the ordering that results: binarizing a read-count matrix conceals a counting artifact and discards the real quantitative signal in the same operation.
What the assay does support
Aggregation. Pseudobulk profiles, metacells, gene activity scores, motif-level summaries and LSI itself all trade locus resolution for signal, and the cell-type-level biology that scATAC-seq has delivered over the past several years is well founded. Peak calling is itself an aggregation step — peaks cannot be called on one cell, so the feature space is defined on pooled data before any cell is examined.5 That has a consequence worth stating explicitly for anyone studying rare populations: a region accessible only in a small subset of cells may never become a column in the matrix, in which case its absence from the results is a property of the peak set rather than a finding about the cells.
| Claim as usually written | Support | What is actually behind it |
|---|---|---|
| “This peak is accessible in this cell type” | SUPPORTED | Aggregation across hundreds or thousands of cells, which is where the assay's information lives. |
| “This cell is closed at this enhancer” | NOT SUPPORTED | More than 75% of peaks carry too little information to distinguish closed from unsampled in one cell. |
| “X% of cells in this cluster are accessible here” | DETECTION RATE | A detection frequency, not an open fraction. It scales with per-cell depth and is not comparable across datasets. |
| “Condition A cells have fewer accessible regions” | DEPTH STATISTIC | Peaks detected per cell is a function of fragments per cell until depth is matched and shown to be matched. |
| “This rare population has a distinct profile” | BOUNDED BY FEATURES | The peak set was called on pooled data; peaks specific to a small population may not be columns at all. |
The lower four rows fail through one mechanism. An absence of observation enters the matrix as a value, and every operation downstream treats it as one.
What to do about it
- Count fragments, not insertion ends. Know which convention produced the matrix you were handed; paired-insertion counting removes an artifact that binarization merely conceals.
- Stop describing TF-IDF as depth normalization. It is a weighting scheme that works well enough for clustering and does not return depth-normalized counts.
- If you drop the first LSI component, say so — and check that it is not carrying biology, since the number of open regions in a cell type is itself informative.
- Report median fragments per cell, TSS enrichment and FRiP alongside any accessibility claim. Without them the reader cannot tell how much of a zero pattern is sampling.
- Never interpret a single-cell, single-locus zero. Not in a figure, not in a supplementary table, not as an intermediate in a downstream calculation.
- Match depth before comparing region counts across conditions. A difference in accessible peaks per cell is a depth difference until demonstrated otherwise.
- Prefer models with an explicit observation probability over transformations that assume the zeros are values.
- Do differential accessibility on aggregates and state the aggregation — which cells, pooled how, at what depth.
- Record how the peak set was built and on which cells. Conclusions about rare populations are bounded by a feature space defined before those populations were identified.
The container carries the interpretation
Most entries in this series turn on a number that means less than it appears to, or a word that quietly upgrades a claim. This one turns on a data structure. A count matrix is not a neutral container; it arrives with a meaning inherited from transcriptomics, where rows are observations, columns are features, and a zero is a measurement of near-absence. Single-cell ATAC-seq fills that container with something else — a thin, biased sample of a molecule present in two copies — and the container's meaning comes along at no charge. Nobody decides to interpret the zero as closed chromatin. The format does it, and every tool built for the format does it again.
None of which is an argument against the assay. scATAC-seq at the cell-type level has been genuinely transformative, the sparsity is quantified rather than suspected, and the field has been unusually candid about it: the sharpest statement of the limit comes from authors who built a model specifically to find out where it sits, and reported that their own model could not clear it. Assay sensitivity is improving. Knowing the ceiling is what makes the data usable — you aggregate deliberately, and design the experiment around a cell-type answer, rather than discovering at review that a single-cell claim was a depth measurement in disguise.
The transposase visited a small and biased sample of the open chromatin in every cell and faithfully reported what it found. Everywhere it did not go, the matrix records a zero. A zero has always looked like an answer.
References
- Kwok AWC, Shim H, McCarthy DJ. A hierarchical, count-based model highlights challenges in scATAC-seq data analysis and points to opportunities to extract finer-resolution information. Genome Biology. 2025;26:282. doi:10.1186/s13059-025-03735-y
- Chen H, Lareau C, Andreani T, et al. Assessment of computational methods for the analysis of single-cell ATAC-seq data. Genome Biology. 2019;20:241. doi:10.1186/s13059-019-1854-5
- Martens LD, Fischer DS, Yépez VA, Theis FJ, Gagneur J. Modeling fragment counts improves single-cell ATAC-seq analysis. Nature Methods. 2024;21(1):28–31. doi:10.1038/s41592-023-02112-6
- Miao Z, Kim J. Uniform quantification of single-nucleus ATAC-seq data with paired-insertion counting (PIC) and a model-based insertion rate estimator. Nature Methods. 2024;21(1):32–36. doi:10.1038/s41592-023-02103-7
- Chen M. Capturing cell-type-specific activities of cis-regulatory elements from peak-based single-cell ATAC-seq. Cell Genomics. 2025;5(3):100806. doi:10.1016/j.xgen.2025.100806
- Li Z, Kuppe C, Ziegler S, et al. Chromatin-accessibility estimation from single-cell ATAC-seq data with scOpen. Nature Communications. 2021;12:6386. doi:10.1038/s41467-021-26530-2
- Luo S, Germain PL, Robinson MD, von Meyenn F. Benchmarking computational methods for single-cell chromatin data analysis. Genome Biology. 2024;25:225. doi:10.1186/s13059-024-03356-x
- Teo AYY, Squair JW, Courtine G, Skinnider MA. Best practices for differential accessibility analysis in single-cell epigenomics. Nature Communications. 2024;15:8805. doi:10.1038/s41467-024-53089-5
- Stuart T, Srivastava A, Madad S, Lareau CA, Satija R. Single-cell chromatin state analysis with Signac. Nature Methods. 2021;18(11):1333–1341. doi:10.1038/s41592-021-01282-5
- Seufert I, Sant P, Bauer K, Syed AP, Rippe K, Mallm JP. Enhancing sensitivity and versatility of Tn5-based single cell omics. Frontiers in Epigenetics and Epigenomics. 2023;1:1245879. doi:10.3389/freae.2023.1245879
Related in this series: Ambient RNA in Single-Cell Data, Normalization Assumptions, Artifact Regions in ChIP-seq and ATAC-seq.

