Cell Biology · Advanced: Modern Techniques

11.5 Genomics, Transcriptomics, and Proteomics

11 min read
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 4 sections
  1. The college version
  2. Eli explains
  3. Study tools
  4. Sources & references

The college version

Core Explanation

The central dogma — DNA → RNA → protein — describes a unidirectional flow of information. But measuring each level reveals distinct, often discordant, views of the cell. Genomics tells us what the cell can do (its blueprint); transcriptomics tells us what it is trying to do (its active instructions); proteomics tells us what it is actually doing (its functional machinery). This topic covers the experimental approaches to each, their relative strengths, and — critically — why they do not always agree.


Genomics: The Blueprint

Technologies

Next-Generation Sequencing (NGS) has largely replaced Sanger sequencing for genome-scale projects. The dominant platform (Illumina) uses sequencing-by-synthesis with reversible dye terminators: fluorescently labeled, 3′-blocked dNTPs are incorporated one base at a time, imaged, and cleaved to allow the next cycle. This generates millions of short reads (50–300 bp) in parallel.

Third-generation sequencing (PacBio SMRT, Oxford Nanopore) produces long reads (10–100+ kb) that span repetitive regions and structural variants inaccessible to short-read platforms.

Applications

ApplicationDescription
Whole-genome sequencing (WGS)Determine the complete DNA sequence. Identifies SNPs, indels, copy-number variants (CNVs), and structural variants (SVs).
Whole-exome sequencing (WES)Sequence only protein-coding regions (~1–2% of the genome). Cost-effective for Mendelian disease gene discovery.
Comparative genomicsCompare genomes across species or populations to identify conserved elements, evolutionary innovations, and disease-associated variants.
MetagenomicsSequence DNA from environmental or microbiome samples to characterize community composition and functional potential.

Key Limitations

  • Static: The genome is (mostly) the same in every cell of an organism throughout its life. Genomics reveals potential, not current activity.
  • Non-coding interpretation: >98% of the human genome is non-coding; assigning function to regulatory elements remains an enormous challenge.

Transcriptomics: The Active Instructions

RNA-Seq Workflow

Step 1: RNA Extraction Total RNA is extracted and assessed for integrity (RNA Integrity Number, RIN). Even modest degradation introduces 3′ bias and reduces detection of low-abundance transcripts.

Step 2: Enrichment/Depletion

  • Poly-A selection: Oligo(dT) beads capture mRNA via poly-A tails. Enriches for mRNA but misses non-polyadenylated transcripts (histone mRNAs, some lncRNAs).
  • rRNA depletion: Ribosomal RNA (>80% of total RNA) is selectively removed, retaining all non-ribosomal species. Preferred for degraded samples (e.g., FFPE tissue) and lncRNA discovery.

Step 3: Library Preparation RNA is fragmented, reverse-transcribed to cDNA, and adapters are ligated. For strand-specific libraries, the second strand is marked (e.g., with dUTP), enabling unambiguous determination of which DNA strand was the original template — critical for distinguishing overlapping sense and antisense transcripts.

Step 4: Sequencing Libraries are sequenced to a depth appropriate for the application: 20–50 million reads/sample for standard differential expression; 50–100 million for transcript assembly and isoform analysis.

Step 5: Alignment and Quantification Reads are aligned to a reference genome (STAR, HISAT2) or pseudoaligned to a transcriptome (Salmon, kallisto). Expression is quantified as counts per gene or transcript, then normalized to account for library size and gene length:

  • TPM (Transcripts Per Million): Normalizes for gene length and library size; values across samples are comparable.
  • FPKM/RPKM: Similar to TPM; TPM is generally preferred for cross-sample comparison.

Step 6: Differential Expression Analysis Tools (DESeq2, edgeR, limma-voom) model count data using a negative binomial distribution and identify genes with statistically significant changes between conditions. Results are typically visualized as:

  • Volcano plot: log₂(fold change) on x-axis vs. −log₁₀(adjusted p-value) on y-axis. Significantly up- and down-regulated genes appear at the extremes.
  • Heatmap: Clustering of samples and genes reveals patterns of co-regulation.

Key Limitation: RNA Abundance ≠ Protein Abundance

This is one of the most important — and frequently ignored — facts in molecular biology. Genome-wide studies consistently find that mRNA levels explain only ~30–40% of the variance in protein abundance (Schwanhäusser et al., 2011; Vogel & Marcotte, 2012). The intervening factors include:

  1. Translation rate: Ribosome occupancy varies across transcripts; some mRNAs are efficiently translated, others are sequestered in P-bodies or stress granules.
  2. Protein half-life: Proteins vary dramatically in stability — from minutes (cyclins, c-Fos) to months (crystallins in the lens). A stable protein at low mRNA can be more abundant than an unstable protein at high mRNA.
  3. Post-translational regulation: Ubiquitination, autophagy, and proteasomal degradation dynamically modulate protein levels independent of synthesis.
  4. Cellular heterogeneity: Single-cell RNA-Seq reveals that mRNA levels vary widely among individual cells; bulk transcriptomics averages this out, masking subpopulations.

Practical rule: Transcriptomics identifies candidate genes. Proteomics (or at minimum, Western blot validation) is required to confirm that changes in mRNA translate to changes in protein.


Proteomics: The Functional Machinery

Proteomics measures the complete complement of proteins — their identities, abundances, modifications, interactions, and localizations. It is fundamentally more challenging than genomics or transcriptomics because:

  1. No amplification: Unlike DNA/RNA, proteins cannot be amplified. Detection must work with whatever is present. This makes low-abundance proteins (transcription factors, signaling kinases) notoriously difficult to detect.
  2. Dynamic range: Protein abundances span >7 orders of magnitude in human cells (albumin in plasma: ~40 mg/mL; cytokines: ~pg/mL). No single method covers this range.
  3. Chemical diversity: The 20 amino acids exhibit vastly different chemical properties (charge, hydrophobicity, size), and post-translational modifications (PTMs) add hundreds of additional forms. No single separation or detection method works for all proteins.
  4. Splice isoforms and proteoforms: One gene can produce multiple protein isoforms via alternative splicing, PTMs, and proteolytic processing. Each proteoform may have distinct function.

Mass-Spectrometry-Based Proteomics Workflow

Step 1: Protein Extraction and Digestion Proteins are extracted under denaturing conditions, reduced (DTT/TCEP) and alkylated (iodoacetamide) to block cysteine residues, and digested with trypsin — a protease that cleaves C-terminally to lysine and arginine, producing peptides of predictable masses ideal for MS analysis.

Step 2: Peptide Separation Peptides are separated by reversed-phase liquid chromatography (LC), typically with a gradient of increasing acetonitrile, which elutes peptides by hydrophobicity. Multidimensional separation (e.g., strong cation exchange + reversed phase) increases coverage.

Step 3: Mass Spectrometry (MS) Eluting peptides are ionized (electrospray ionization, ESI) and introduced into the mass spectrometer. Two stages:

  • MS1 (survey scan): Measures the mass-to-charge (m/z) ratios of intact peptide ions. The most abundant ions are selected for fragmentation.
  • MS2 (tandem MS / MS/MS): Selected precursor ions are fragmented (collision-induced dissociation, CID) and the fragment masses are recorded. The fragmentation pattern serves as a peptide "fingerprint."

Step 4: Peptide Identification MS/MS spectra are searched against a protein sequence database using algorithms (Mascot, SEQUEST, MaxQuant/Andromeda) that perform in silico digestion and match experimental spectra to theoretical fragmentation patterns. False discovery rate (FDR) is controlled by searching against a decoy (reversed) database; a 1% FDR at the peptide level is standard.

Step 5: Quantification

MethodPrincipleNotes
Label-free quantification (LFQ)Compare MS1 peak intensities across runsSimple but susceptible to run-to-run variation
Stable isotope labeling (SILAC, iTRAQ, TMT)Samples are labeled with isotopes and multiplexed in a single runReduces technical variation; TMT allows up to 18-plex
Data-independent acquisition (DIA/SWATH)Fragment all peptides in defined m/z windows; quantify from fragment ionsReproducible; suitable for large cohorts

Post-Translational Modification (PTM) Analysis

PTMs (phosphorylation, ubiquitination, acetylation, glycosylation, etc.) are identified by characteristic mass shifts in MS/MS spectra. Enrichment strategies (e.g., immobilized metal affinity chromatography for phosphopeptides) are essential because modified peptides are often substoichiometric and invisible in unfractionated samples. Phosphoproteomics routinely identifies >10,000 phosphorylation sites, but assigning function to individual sites remains a major bottleneck.


Comparing the Three Omics

FeatureGenomicsTranscriptomicsProteomics
What is measuredDNA sequence and structural variationRNA transcript abundanceProtein abundance, modifications, interactions
Core technologyNGS (Illumina, long-read)RNA-SeqLC-MS/MS
Amplification possible?Yes (PCR)Yes (PCR via cDNA)No
Dynamic rangeLow (binary per locus, or copy number)Moderate (~4–5 orders)Extreme (>7 orders)
InformationBlueprint — what the genome can encodeInstructions — what the cell is trying to expressFunctional state — what the cell is actually doing
Response timeNone (static)Minutes to hoursVariable (minutes to days, depending on protein half-life)
Typical coverage>30× genome-wide>10,000 genes3,000–8,000 proteins (unfractionated cell lysate)
Splicing/isoformsDetectable (gene structure)Pervasive; isoform-level quantification challengingProteoforms visible but challenging to resolve

Readout

  • Genomics: VCF files (variants), BAM files (alignments), genome browsers (IGV).
  • Transcriptomics: Count matrices, volcano plots, MA plots, PCA plots (quality control), heatmaps, gene set enrichment analyses.
  • Proteomics: Peptide-spectrum matches, protein groups tables, volcano plots, PTM occupancy ratios.

Strengths

  • Integration: Combined multi-omics provides a systems-level view unattainable from any single layer.
  • Discovery-based: Unlike hypothesis-driven methods (qPCR, Western), omics approaches are unbiased and can reveal unexpected players.
  • Quantitative: All three can be quantitative, enabling statistical comparison across conditions.

Common Interpretation Errors

  • "My RNA-Seq shows Gene X is upregulated 5-fold, so the protein must be upregulated too." mRNA-protein correlation is modest. A 5-fold mRNA change may produce no protein change, or vice versa. Transcriptomic results are hypotheses; proteomic validation is evidence.
  • "The most differentially expressed gene is the most important." Statistical significance (p-value) is driven by effect size AND variance. A gene with a 1.5-fold change and low variance can be more "significant" than a 10-fold change with high variance. Biological importance and statistical significance are not the same.
  • "My proteomics experiment identified 3,000 proteins, so I'm missing the other ~7,000." Coverage depends on sample complexity, fractionation, and instrument sensitivity. Missing proteins are not necessarily absent — they may be below the detection limit, poorly ionized, or obscured by highly abundant proteins. A negative proteomics result is rarely proof of absence.
  • "The volcano plot threshold (log₂FC > 1, p < 0.05) defines all the 'real' hits." Arbitrary cutoffs discard biologically meaningful changes with small effect sizes (e.g., transcription factors where 2-fold is biologically significant) and may include statistically significant but biologically trivial changes.

Quick Questions

Q1: You perform RNA-Seq on cells treated with a drug and find that Gene W is upregulated 8-fold (adjusted p = 10⁻¹²). You run a Western blot and see no change in Gene W protein. List three biologically plausible explanations.

Answer
  1. Translational repression: The increased mRNA is not efficiently translated — e.g., miRNA-mediated repression, or the mRNA is sequestered in stress granules or P-bodies.
  2. Protein degradation: Protein half-life is short, and degradation rates have increased to match the increased synthesis, resulting in no net change in steady-state protein levels.
  3. Kinetics: You harvested at a time point where mRNA has risen but protein has not yet accumulated (translation and protein accumulation lag behind transcription, sometimes by hours).
  4. Isoform mismatch: Your Western blot antibody recognizes an epitope present in only one isoform, while RNA-Seq counts all isoforms collectively.

Q2: Why is proteomics coverage typically lower than transcriptomics coverage, and what strategies can improve it?

Answer

Proteomics cannot amplify proteins; it must detect what is present. Abundant proteins (ribosomal proteins, cytoskeletal proteins, metabolic enzymes) dominate the signal and can mask low-abundance species. Strategies to improve coverage:

  • Fractionation: Separate peptides or proteins into multiple fractions before MS (e.g., SCX, high-pH reversed-phase), reducing sample complexity per run.
  • Depletion: Remove highly abundant proteins (e.g., albumin/IgG from plasma; ribosome depletion from cell lysates) to unmask low-abundance species.
  • Longer gradients: Longer LC gradients increase peptide separation and MS sampling depth.
  • Gas-phase fractionation: Acquire multiple runs covering different m/z ranges.
  • Instrument sensitivity: Newer instruments (Orbitrap Astral, timsTOF) achieve deeper coverage in shorter times.

Q3: What does a volcano plot show, and why might a gene with a 10-fold change not appear as a "hit" while a gene with a 1.5-fold change does?

Answer

A volcano plot displays log₂(fold change) on the x-axis and −log₁₀(adjusted p-value) on the y-axis. Each point is a gene. "Hits" are typically defined as genes exceeding thresholds on both axes (e.g., |log₂FC| > 1 and adjusted p < 0.05). A gene with a large fold change (10-fold, log₂FC ≈ 3.3) can fail significance if it has very high variance across biological replicates (e.g., noisy expression in one replicate inflates the standard error, reducing the test statistic). Conversely, a gene with a small fold change (1.5-fold, log₂FC ≈ 0.58) can be highly significant if it is measured with extreme precision (low variance across replicates, high read counts). Significance measures confidence that the change is real, not magnitude of the change.


Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

Think of a cell as a city. Genomics is the city's master blueprint — it shows what could potentially be built on every lot, but most of it is just plans. Transcriptomics is like reading all the building permits currently active — it tells you what construction is being attempted. Proteomics is like walking the streets and looking at what's actually standing — the real buildings, not the permits. And just because there's a building permit (mRNA) for a skyscraper doesn't mean the skyscraper (protein) actually got built — maybe the construction crew was slow, or the permit expired, or the city inspector (the cell's quality control) tore it down.


Keep learning

Ready to build on this? Continue to the next lesson.

Study tools & related lessonsYou’ll learn to · Related

You’ll learn to

  • By the end of this section, you should be able to:
  • Compare and contrast genomics, transcriptomics, and proteomics in terms of what each measures, the core technology, and the biological questions each addresses.
  • Describe the workflow of an RNA-Seq experiment from RNA extraction through differential expression analysis.
  • Explain why RNA abundance does not equal protein abundance and identify the biological factors that intervene.
  • Outline the principles of mass-spectrometry-based proteomics, including peptide identification and quantification.
  • Discuss the unique challenges of proteomics compared to genomics and transcriptomics.
  • Interpret a volcano plot of differential expression results.
  • ---

Sources & references

  1. NCBI Bookshelf: Brown, T. A. *Genomes* (2nd ed.), Chapter on Transcriptomes and Proteomes. Available at

This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.