Transcript Abundance and Protein Abundance: What an RNA-Seq Result Does Not Establish

Transcript Abundance and Protein Abundance: What an RNA-Seq Result Does Not Establish | Zetobit
Zetobit · Bioinformatics Insight Series
Transcript Abundance and Protein Abundance Header card. Two vertical columns of points represent transcript measurements on the left and protein measurements on the right, joined by connecting lines of markedly uneven thickness, illustrating that the relationship between the two layers differs greatly from gene to gene. TRANSCRIPT PROTEIN ZETOBIT BIOINFORMATICS INSIGHT SERIES Transcript Abundance and Protein Abundance What an RNA-seq result does not establish 0.43 median mRNA–protein correlation, 13 studies, standardized pipeline

Multi-Omics Integration

Transcript Abundance and Protein Abundance: What an RNA-Seq Result Does Not Establish

Differential expression is routinely read as a statement about protein. Across thirteen large proteogenomic studies reanalyzed under a single pipeline, the reported median mRNA–protein correlation is 0.43 — and a substantial share of that shortfall turns out to be measurement error rather than biology. Both halves of that finding matter, and they point in opposite directions.

An RNA-seq experiment ends with a list of genes and fold changes. The discussion section that follows almost always talks about proteins: a pathway is activated, an enzyme is upregulated, a receptor is lost. That translation from one noun to the other is so routine it rarely gets stated as an assumption, and it is usually approximately right. mRNA is the template for protein synthesis, and the gene expression pathway is hierarchical — a transcriptionally silent gene cannot be upregulated by stabilizing a protein that is not being made.1

But "approximately right" carries a distribution, and the shape of that distribution is not what most people assume. The interesting part is not that the correlation is imperfect. It is that the imperfection has two entirely different sources — real post-transcriptional regulation and ordinary measurement noise — and that the field spent years attributing the second to the first.

Two questions that are both called "correlation"

Before any number means anything, the question it answers has to be pinned down. There are two, and they are not interchangeable.1

Across-gene correlation asks: within one sample, are the most abundant transcripts also the most abundant proteins? This compares absolute abundances across thousands of genes under a single condition. Typical Pearson coefficients for mammalian tissues run around 0.6, meaning roughly 40% of the variability in protein levels tracks with variability in mRNA levels.1

Within-gene correlation asks: across many samples, do the samples with the most transcript for gene X also have the most protein X? This is the question a differential expression analysis is actually asking, and it is the one that matters for nearly every applied use — comparing tumor to normal, treated to untreated, responder to non-responder.

These are different measurements answering different questions, and conflating them produces confident nonsense in both directions. A strong across-gene correlation does not license the inference that a two-fold transcript change implies a two-fold protein change. Median per-gene protein-to-mRNA ratios can produce reasonable order-of-magnitude estimates of absolute protein abundance across genes — but this works partly because the dynamic range of protein copy number across genes is enormous, far larger than the range over which any single gene's abundance varies across tissues.1 Predicting that ribosomal proteins are more abundant than transcription factors is not the same skill as predicting which tumor has more of a given protein. Notably, the same ratio-based prediction works even when the ratios are computed from randomly shuffled mRNA levels — a result that should discipline how much the across-gene success is taken to mean.2

The across-gene question and the within-gene question share a name, a statistic, and almost nothing else.

What the within-gene numbers actually are

Reported correlations vary widely across proteogenomic studies, but much of that spread is an artifact of inconsistent methodology — different studies used mean versus median, Pearson versus Spearman, and wildly different protein inclusion criteria. Upadhya and Ryan recalculated all of them through a single standardized pipeline, restricting to proteins measured in at least 80% of samples.3

Within-gene mRNA–protein correlation across proteogenomic studies, as originally reported versus recalculated under a standardized pipeline. Recomputed values are median Spearman.3
Dataset Year As reported Recomputed
Lung adenocarcinoma20200.530.55
Head & neck squamous cell20210.520.54
GTEx, 32 healthy tissues20200.460.51
Glioblastoma20210.50
Endometrial cancer20200.480.48
Cancer Cell Line Encyclopedia20200.480.46
Breast cancer20200.410.44
Breast cancer20160.390.42
Ovarian cancer20160.450.41
Clear cell renal carcinoma20190.430.41
NCI-60 cell lines20190.36
Colon cancer20190.480.27
Colon and rectal cancer20140.230.21

Across all thirteen studies the authors report a median recalculated correlation of 0.43, with a maximum of 0.55 and a minimum of 0.21.3 Two features of this table deserve attention beyond the headline number.

First, the colon cancer row: reported at 0.48, recomputed at 0.27. The original figure was calculated on only the 10% most variable proteins, which have higher-than-average correlation.3 Nothing was misrepresented — the inclusion criterion was stated — but the number that entered general circulation was not comparable to the numbers it was being compared against. This is a recurring hazard in any field where a single summary statistic gets quoted downstream of its methods section.

Second, the trend over time. Studies published after 2019 average 0.49; those from 2016 or earlier average 0.35. This is not explained by which cancers were studied — the two tumor types profiled twice both improved.3 Something about the measurement got better. That observation is the thread worth pulling.

How much of the gap is noise

If a moderate correlation reflects genuine post-transcriptional regulation, it is a biological finding. If it reflects the fact that mass spectrometry quantifies some proteins poorly, it is an instrument limitation wearing a biology costume. Distinguishing these requires replicate measurements — the same sample, profiled twice.

Three proteogenomic studies had them. The result: median protein–protein reproducibility across replicates was 0.72 in the Cancer Cell Line Encyclopedia data (same lab, one year apart), 0.57 for ovarian tumors (two different labs), and 0.28 for colon tumors (two different mass spectrometry techniques).3 Replicate measurements of the same protein, where no post-transcriptional regulation can possibly intervene, agree only moderately. And critically, the per-protein reproducibility ranged from near zero to 1.0 — some proteins are measured well, others are barely measured at all.

Stratifying proteins into deciles by reproducibility and computing mRNA–protein correlation within each decile gives the decisive result: correlation rises monotonically with reproducibility. The gap between the least and most reproducible deciles was 0.33 to 0.37 depending on the study.3 Protein measurement reproducibility alone explains roughly 14–23% of the variance in mRNA–protein correlation. Transcript reproducibility contributes independently, with a median of about 15% variance explained across studies.3

Relationship between protein measurement reproducibility and observed mRNA-protein correlation A schematic showing that proteins binned by how reproducibly they can be measured show systematically increasing mRNA-protein correlation, with the least reproducible decile far below the study median and the most reproducible decile far above it. 0.1 0.3 0.5 0.7 study median correlation lowest decile highest decile Δ ≈ 0.33–0.37 Proteins binned by measurement reproducibility across replicates → mRNA–protein correlation The correlation you observe depends on how well you measured
Schematic of the observed relationship. Proteins that replicate measurements agree on show substantially higher mRNA–protein correlation than proteins they do not, across every study examined — in tumors and healthy tissue, and under both data-dependent and data-independent acquisition. Bar heights are illustrative of the reported trend, not exact decile values.

The consequence is uncomfortable for a body of published interpretation. Pathway enrichment analyses have repeatedly found that metabolic pathways show higher-than-average mRNA–protein correlation, and this has been read as evidence that those pathways are less post-transcriptionally regulated. After adjusting for measurement reproducibility, that enrichment disappears — the metabolic proteins were simply measured more reproducibly.3 The ribosomal and housekeeping-complex signal in the opposite direction survived the adjustment. One inference was real biology; the other was an instrument artifact that had been given a mechanistic story.

The biology that remains

Correcting for noise does not dissolve post-transcriptional regulation. It relocates the evidence for it.

Protein-level buffering

The most robust evidence comes from cases where mRNA is forced to change and protein declines to follow. In tumors, copy-number variation drives transcript abundance almost proportionally — but 23–33% of proteins are post-transcriptionally buffered against that change, with strong enrichment for members of protein complexes.4 The mechanism is well characterized: unassembled subunits of multiprotein complexes are recognized as orphans and degraded, so a complex settles at the level of its stoichiometrically limiting member regardless of how much transcript the other subunits have.1 Related findings hold across germline variation: only about a third of mRNA QTLs in human lymphoblastoid lines are associated with protein-level changes.1

This is the single most important practical consequence for cancer genomics. An amplified region produces elevated transcript for everything in it. Which of those genes actually produces elevated protein — and is therefore a plausible driver or target — is not answerable from RNA-seq.

Temporal offset

Any transcriptional change is followed by a delayed protein change, because the system takes time to reach a new steady state. Correlations computed at a single timepoint during a transition can be uninformative not because the relationship is weak but because the protein response has not happened yet.1 In Drosophila development, mRNA levels at one timepoint correlate better with protein levels at a later timepoint than with the matched one.1 A design that samples RNA and protein at the same moment during a dynamic process has built in a mismatch that no analysis can remove.

Spatial disconnection

Correlation presupposes that transcript and protein occupy the same sample. Secreted proteins violate this by construction — plasma and urine contain proteomes assembled from many tissues and almost no corresponding mRNA. Highly polarized cells such as neurons can violate it internally, where the distance between cell body and axon terminal is large enough to decouple the two measurements.1 Any tissue with substantial extracellular matrix or mixed cellular composition carries some version of this problem.

Proliferation state

A dividing cell loses half its biomass every cycle and must synthesize protein in proportion to abundance simply to hold levels constant — which requires transcript. A quiescent cell need only replace what degrades, and many proteins in quiescent cells have half-lives exceeding twenty days.1 Proliferation status is therefore a determinant of how tightly the two layers couple, which means it is a confounder whenever samples differ in proliferative index. Tumor versus normal is exactly such a comparison.

What this changes in practice

The claim an RNA-seq result supports

A differential expression result establishes that transcript abundance differs between conditions. It is evidence that protein abundance may differ, with a per-gene reliability that varies from near-total to near-zero and is partly a property of the gene and partly a property of how well that gene's protein can be quantified. It is not a protein measurement, and the distance between the two is not a constant that can be absorbed into a caveat sentence.

  • Name which correlation is being cited. A quoted figure is meaningless without knowing whether it is across-gene or within-gene, Pearson or Spearman, mean or median, and what the protein inclusion threshold was. The colon cancer row above moved 0.21 on inclusion criteria alone.
  • Treat per-gene reliability as a gene-level property, not a global discount. Complex subunits, ribosomal proteins, and low-abundance proteins with few unique peptides sit at one end; abundant, variable, peptide-rich proteins sit at the other. A conclusion about a specific gene inherits that gene's position, not the study median.
  • Do not attribute low correlation to post-transcriptional regulation without checking measurement reproducibility first. This is the specific error the reproducibility analysis identified in published pathway interpretations. Where replicates exist, use them; where they do not, aggregate reproducibility ranks are published and can be joined in.
  • Watch for genes with no variance. A tightly regulated transcript that does not vary across samples cannot correlate with anything. Low correlation here means low signal, not post-transcriptional control — and the distinction is invisible in the coefficient.
  • Design around the temporal offset. In any perturbation or developmental series, matched-timepoint RNA and protein sampling embeds a lag. Either sample protein later, or model the offset explicitly, or state that the design cannot resolve it.
  • In copy-number contexts, expect buffering. Roughly a quarter to a third of proteins do not follow their amplified transcript. Prioritizing candidate drivers from an amplicon on transcript evidence alone will include genes whose protein never moved.
  • Prefer proteins that can actually be measured for anything downstream. If a marker is destined for a diagnostic panel or a validation assay, reproducibility of quantification is a selection criterion in its own right, independent of biological interest.

Non-redundant, not interchangeable

The framing that serves best is not that one layer is a flawed proxy for the other. It is that they are non-redundant readouts of different stages of the same pathway.1 Transcript data sits closer to the genome and more directly reflects transcription factor activity, chromatin state, and copy number — RNA-seq can even be used to recover copy-number structure. Protein data sits closer to phenotype, is more robust to functionally irrelevant transcript variation, and captures regulation that has no transcript signature at all.

The presence of mRNA means the template for protein synthesis was available. The presence of protein means a transcript existed at the time and place the protein was made. Neither statement is the other, and the analytical work is in being precise about which one the data supports.

mRNA levels evolved to serve future protein synthesis needs — not as a summary of what the cell currently contains.

References

  1. Buccitelli C, Selbach M. mRNAs, proteins and the emerging principles of gene expression control. Nature Reviews Genetics 21:630–644 (2020). doi:10.1038/s41576-020-0258-4
  2. Fortelny N, Overall CM, Pavlidis P, Freue GVC. Can we predict protein from mRNA levels? Nature 547:E19–E20 (2017). doi:10.1038/nature22293
  3. Upadhya SR, Ryan CJ. Experimental reproducibility limits the correlation between mRNA and protein abundances in tumor proteomic profiles. Cell Reports Methods 2(9):100288 (2022). doi:10.1016/j.crmeth.2022.100288
  4. Gonçalves E, Fragoulis A, Garcia-Alonso L, Cramer T, Saez-Rodriguez J, Beltrao P. Widespread post-transcriptional attenuation of genomic copy-number variation in cancer. Cell Systems 5(4):386–398 (2017). doi:10.1016/j.cels.2017.08.013
  5. Edfors F, et al. Gene-specific correlation of RNA and protein levels in human cells and tissues. Molecular Systems Biology 12:883 (2016). doi:10.15252/msb.20167144
  6. Wang D, et al. A deep proteome and transcriptome abundance atlas of 29 healthy human tissues. Molecular Systems Biology 15:e8503 (2019). doi:10.15252/msb.20188503
  7. Vogel C, Marcotte EM. Insights into the regulation of protein abundance from proteomic and transcriptomic analyses. Nature Reviews Genetics 13:227–232 (2012). doi:10.1038/nrg3185
  8. Csárdi G, Franks A, Choi DS, Airoldi EM, Drummond DA. Accounting for experimental noise reveals that mRNA levels, amplified by post-transcriptional processes, largely determine steady-state protein levels in yeast. PLOS Genetics 11(5):e1005206 (2015). doi:10.1371/journal.pgen.1005206

Zetobit is a bioinformatics consulting firm working with biotech, pharma, and academic groups on multi-omics analysis, clinical NGS pipelines, and variant interpretation. The Bioinformatics Insight Series examines where analytical methods and their assumptions diverge.

Previous
Previous

Tumor Purity and Ploidy: Why the Number Every Somatic Result Depends On Has More Than One Valid Answer

Next
Next

False Positives in Metagenomic Classification: Why a Species in the Report Is Not Evidence It Was in the Sample