Normalization Assumptions
Normalization Assumptions
Why a Global Shift Is the One Thing the Method Cannot See
An RNA-seq count is not a measurement of expression. It is a measurement of expression multiplied by an unknown constant — how much material went in, how deeply it was sequenced, how efficiently the library was built. Normalization estimates that constant so counts can be compared, and to estimate it from the data you have to assume something is invariant across samples.
That assumption is the subject here. It is not a tuning parameter and it is not a tool preference: it is a biological claim, made once, that the analysis then depends on entirely.
A note on scope, because this sits next to something already covered. An earlier piece dealt with the choice between DESeq2, edgeR and limma-voom — the dispersion model and the test. This is a step earlier, and the reason it is a different subject is that the tool choice does not touch it. DESeq2's median-of-ratios and edgeR's TMM are different estimators of the same quantity resting on the same premise, and limma-voom conventionally uses TMM outright. When the premise fails, all three fail together, in the same direction, by roughly the same amount. Switching tools is not a remedy because there is nothing to switch to.
What the estimators actually assume
TMM computes a scale factor from a trimmed mean of log-ratios between a sample and a reference, discarding genes with extreme ratios or extreme abundance1. Median-of-ratios takes, for each sample, the median ratio of each gene's count to a geometric-mean pseudo-reference across samples2. Both are robust in the statistical sense: a minority of strongly changed genes will not move them.
Robustness to a minority is exactly the property that becomes a liability against a majority. The premise both rest on is that most genes are not differentially expressed, or that up- and down-regulation are roughly balanced so the middle of the distribution stays put. A review organising normalization methods by their assumptions puts the general form plainly: nearly all existing procedures assume most genes are not differentially expressed and that those which are are equally likely to be up- or down-regulated3.
Under a global shift, that premise fails in every way at once. The same review is precise about it: global up-regulation necessarily produces different amounts of mRNA per cell, which breaks library-size normalization; highly asymmetric expression, which breaks distribution-based methods; and an absence of non-differentially-expressed genes, which breaks housekeeping-gene normalization — and none of these can be rescued without external controls3.
The experiment that made it concrete
The demonstration everyone cites used P493-6 cells, a B-cell line carrying a tetracycline-repressible MYC transgene, so c-Myc can be titrated in an otherwise identical background. Cells with high c-Myc produce two- to three-fold higher levels of the same RNA species found in low-Myc cells4.
Equal numbers of high- and low-Myc cells were counted on a haemocytometer and harvested, spike-in standards were added on a per-cell basis, and the RNA was sequenced. Processed by standard normalization, the results suggested some genes were unchanged while others increased or decreased. Normalized instead against the spike-ins — which reflect cell number rather than RNA mass — transcript levels increased for the vast majority of genes4.
An independent group using whole Drosophila cells as spike-ins rather than synthetic transcripts reproduced the effect and added the number that makes it tangible: total RNA extracted from the same number of cells differed by nearly 1.8-fold between high- and low-MYC samples, and while spike-in calibration showed global induction, endogenous normalization of the same data gave log fold changes scattered around zero — implying similar numbers of induced and repressed genes5.
Why you cannot detect this from the data
Here is the part that makes this different from an ordinary assumption violation, and it is the reason the problem persists despite being well documented.
Most broken assumptions leave a residue. A misspecified dispersion model produces poorly calibrated p-values you can check against a null. A batch effect shows up in a PCA. This one does not, because the normalization step enforces the assumption on its output. If you scale each sample so that the bulk of genes has a log-ratio near zero, then the bulk of genes will have a log-ratio near zero — whatever was true of the cells.
So the diagnostic you would naturally reach for is the artifact. A volcano plot with comparable numbers of up and down genes is what a well-behaved experiment looks like, and it is also precisely what a globally amplified experiment looks like after normalization. The violation erases its own evidence.
The mechanism behind the inverted directions is worth stating without the statistics. Sequencing a library samples a fixed number of reads from whatever was loaded, so a read count is a share — a gene's fraction of the material, not its quantity. If every gene in the cell triples, every share stays where it was. If most genes triple and one holds steady, that one gene's share falls, and the pipeline reports it as down-regulated. It is down as a proportion of the transcriptome, which is true and is not what the reader takes from the phrase.
Where this is a live risk
It is worth being clear that this is not a general indictment of TMM or median-of-ratios. For most experiments — two tissues, a modest treatment effect, a disease-versus-control contrast where a few hundred genes move — the assumption is close enough to true, and these estimators are the right tools, well validated and appropriately robust.
The risk is concentrated in a recognisable set of situations, and the common feature is that the perturbation acts on transcription globally rather than on particular genes.
| Situation | Why the premise is at risk |
|---|---|
| MYC or MYCN amplification | High MYC and MYCN levels substantially increase overall mRNA per cell6; a comparison across MYC status is a comparison across total RNA content |
| Drugs targeting the transcription machinery | Broad inhibition changes total output rather than redistributing it, so treated and control cells hold different amounts of RNA |
| Cell activation and proliferation states | Resting and activated cells differ in biosynthetic capacity, so RNA per cell differs before any specific programme is considered |
| Global shutdown — stress, apoptosis, host shutoff | The same problem with the sign reversed; unchanged genes appear up-regulated |
| Comparisons across cell size or ploidy | Content scales with the cell, and per-cell and per-microgram questions diverge |
These are prompts to check, not verdicts. The point is that whether the assumption holds is a question about the biology of the perturbation, answerable before sequencing and not afterwards from the counts.
The fix, and what it costs
The only way to recover an absolute scale is to add something whose quantity you control. That means spike-ins, added on a per-cell basis rather than per microgram of RNA — the distinction is the whole point, since adding a fixed amount of spike-in to a fixed mass of RNA re-imposes the equal-total-RNA assumption you were trying to escape.
Two honest caveats keep this from being a clean recommendation.
First, spike-ins have their own assumptions, and the review that lays out the taxonomy is direct about it: spike-in normalization relies on controls that should have the same expression across conditions, these methods come with their own assumptions, and it is not clear those assumptions can always be trusted3. Synthetic ERCC transcripts differ from endogenous mRNA in length, structure and poly-A characteristics, and their behaviour across experiments has been questioned since the earliest assessments3. Whole-cell spike-ins from a different species are one response to that, since the added material is real mRNA in real cells5.
Second, spike-ins require a decision at the bench, before the RNA is extracted. Nothing computational recovers the information afterwards. For a public dataset, or an experiment already run, the absolute question is closed and the honest response is to state that the analysis is relative and to avoid claims that depend on total output.
An accurate cell count with matched input is a partial substitute worth knowing about — it was the basis of the P493-6 design4 — but it requires countable cells and does not survive tissue dissociation or clinical specimens.
The same assumption, one level down
Single-cell RNA-seq inherits this problem in a form that is easier to overlook, because there the scale factor is usually the cell's own total count. Dividing each cell by its library size and multiplying by a constant encodes the same premise per cell: that differences in total counts are technical. Where cell types genuinely differ in RNA content — a plasma cell against a resting lymphocyte, a neuron against a glial cell — part of what is being divided out is biology.
The consequence is the familiar one, relocated: a cell type with high RNA content will show its genes flattened toward the population mean, and comparisons of expression between cell types of different size carry the same directional risk as the bulk case. The remedies are the same in structure and harder in practice, which is why the honest move in most single-cell work is to treat expression values as compositional and say so, rather than to claim per-cell quantities the protocol never established.
What to do about it
- Decide before sequencing whether your perturbation could change total RNA per cell. This is a biology question with a cheap answer and no post-hoc equivalent.
- Where the answer is yes or unknown and the claim matters, add per-cell spike-ins. Per-cell, not per-microgram — the second re-imposes the assumption.
- Do not read a balanced volcano plot as reassurance. It is the shape normalization produces regardless.
- Do not treat the choice of DE package as relevant here. Median-of-ratios and TMM share the premise12, so agreement between them is not corroboration.
- Record total RNA yield per cell or per unit input alongside the counts. A large difference between arms is the cheapest available warning, and it is usually already measured and then discarded.
- Phrase conclusions to match the scale you have. “Enriched relative to the rest of the transcriptome” is supported; “up-regulated” implies a per-cell claim that a relative measurement does not make.
- Be especially careful with down-regulation in a globally activating system. The apparently repressed genes are the ones a global amplification manufactures, and they are the most likely to be an artifact of the reference having moved.
- Report the normalization method as a stated assumption, not a step — one methods sentence naming the method and why its premise is plausible for the contrast.
- When reanalysing public data across a global perturbation, say what cannot be answered. The information was not collected; no amount of reprocessing supplies it.
The shape of the error
This series keeps returning to results that are correct as computed and wrong as read. This one has an unusually clean version of that structure, because the arithmetic is not merely defensible — it is exactly right, given what it was asked.
Normalization was asked: what scale factor makes these samples comparable, assuming most genes did not change? It answered correctly. The answer is wrong only because the assumption was false, and the assumption was not something the method could evaluate — it was supplied by whoever ran the pipeline, usually by not thinking about it, because it is the default and the default is right most of the time.
What makes it worth an article rather than a footnote is the direction of the failure. Not noisier estimates, not lower power: fold changes that point the wrong way. A gene whose transcription did not change at all is reported as significantly down-regulated, with a small p-value and a tight confidence interval, in a system where everything else went up.
The counts are real. The scale factor is a hypothesis, and it is the one number in the analysis that nothing downstream can check.
References
- Robinson MD, Oshlack A. A scaling normalization method for differential expression analysis of RNA-seq data. Genome Biology 2010;11:R25. genomebiology.biomedcentral.com — gb-2010-11-3-r25
- Anders S, Huber W. Differential expression analysis for sequence count data. Genome Biology 2010;11:R106. genomebiology.biomedcentral.com — gb-2010-11-10-r106
- Evans C, Hardin J, Stoebel DM. Selecting between-sample RNA-Seq normalization methods from the perspective of their assumptions. Briefings in Bioinformatics 2018;19(5):776–792. academic.oup.com/bib/19/5/776
- Lovén J, Orlando DA, Sigova AA, et al. Revisiting global gene expression analysis. Cell 2012;151(3):476–482. cell.com — S0092-8674(12)01226-3
- External calibration with Drosophila whole-cell spike-ins delivers absolute mRNA fold changes from human RNA-Seq and qPCR data. BioTechniques 2017. Author list not captured — verify before publication. tandfonline.com — 10.2144/000114514
- Reported key features of MYC and MYCN driven transcriptional amplification, including the finding that high MYC and MYCN levels substantially increase overall mRNA expression per cell and that per-cell spike-in RNA-seq more accurately reflects global expression change. Secondary source; trace to the primary neuroblastoma study before publication. researchgate.net — Revisiting Global Gene Expression Analysis

