Cell Biology · Modern Techniques

Proteomics

5 min read
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 7 sections
  1. In 30 seconds
  2. Why this matters
  3. The college version
  4. Eli explains
  5. Key takeaway
  6. Study tools
  7. Sources & references

In 30 seconds

Proteomics is the large-scale study of the proteome — the complete set of proteins expressed by a cell, tissue, or organism, including their abundances, modifications, and interactions. Because protein levels are governed by translation, turnover, and post-translational modification (not just mRNA), the proteome must be measured directly. The core method is mass spectrometry (MS): proteins are digested into peptides, ionized, and separated by their mass-to-charge ratio (m/z); selected peptides are fragmented (MS/MS), and the fragment spectra are matched against a database to identify the parent proteins and quantify their abundance. Proteomics measures protein identity, amount, and modification — the closest large-scale readout to the functional output of the genome.

Why this matters

Proteomics closes the gap between genome/transcriptome and phenotype. It identifies disease biomarkers in blood, reveals drug targets and drug mechanisms, maps signaling networks through phosphoproteomics, and characterizes protein complexes and modifications that genomics cannot predict. Because most drugs act on proteins, the proteome is the most direct "parts that matter" list for medicine.

The college version

Core Concept

Proteomics is the large-scale study of the proteome — the complete set of proteins expressed by a cell, tissue, or organism, including their abundances, modifications, and interactions. Because protein levels are governed by translation, turnover, and post-translational modification (not just mRNA), the proteome must be measured directly. The core method is mass spectrometry (MS): proteins are digested into peptides, ionized, and separated by their mass-to-charge ratio (m/z); selected peptides are fragmented (MS/MS), and the fragment spectra are matched against a database to identify the parent proteins and quantify their abundance. Proteomics measures protein identity, amount, and modification — the closest large-scale readout to the functional output of the genome.

Key Components

Protein extraction and digestion

  • Proteins are extracted, denatured, and cleaved with a specific protease (usually trypsin, which cuts after lysine/arginine) into predictable peptides.

Liquid chromatography (LC)

  • Peptides are separated by hydrophobicity before entering the mass spectrometer, reducing complexity and increasing coverage.

Mass spectrometry (MS and MS/MS)

  • An ion source ionizes peptides (electrospray); the first analyzer measures m/z (MS1); selected precursor ions are fragmented (collision-induced dissociation) and their fragments measured (MS2), generating a "fingerprint" of each peptide.

Database search and quantification

  • Fragment spectra are matched to theoretical spectra from a protein database (e.g., UniProt) to identify peptides/proteins; abundance is estimated from peptide ion intensities (label-free, or with isobaric tags such as TMT/iTRAQ, or stable isotope labeling).

Post-translational modification (PTM) analysis

  • Mass shifts in fragment spectra reveal modifications (phosphorylation +80 Da, ubiquitination, acetylation, etc.).

Mechanism

  1. Extract and digest. Proteins are isolated and cut with trypsin into peptides.
  2. Separate. Peptides are resolved by liquid chromatography.
  3. Ionize and measure. Peptides are ionized; MS1 records each peptide's m/z.
  4. Fragment. Selected peptides are fragmented; MS2 records the fragment-ion spectrum.
  5. Identify. Software matches MS2 spectra to a database, naming the peptides and their parent proteins.
  6. Quantify. Peptide intensities (or reporter tags) yield relative/absolute protein abundance; PTMs are read from mass shifts.

Energy and Directionality

Mass spectrometry is a physical separation, not an enzymatic reaction: gas-phase ions are accelerated and steered by electric/magnetic fields, and the "signal" is the m/z of each ion. Ionization and fragmentation require input energy (electrospray voltage; collision energy in MS/MS) to break peptide bonds at predictable positions. The information flows directionally — intact peptide mass (MS1) → fragment masses (MS2) → matched sequence — and the biological conclusion (which protein, how much, what modification) is inferred from that physical readout.

Experimental Evidence

  • What it measures: protein identity, abundance, and post-translational modifications across a sample (the proteome).
  • Principle: protease digestion → LC separation → MS1/MS2 mass-to-charge measurement → database matching.
  • Input: protein lysate (cells/tissue); Output: lists of identified proteins with (relative) abundances and detected PTMs.
  • What it can prove: which proteins are actually present and in what relative amounts (unlike mRNA); that a modification (e.g., phosphorylation) occurs at a specific site; protein-complex composition (via affinity purification + MS); differential protein abundance between conditions.
  • What it cannot prove: protein activity or functional state directly (a protein can be abundant but inactive); spatial localization; dynamic transient interactions without specialized approaches; and it may miss low-abundance or poorly ionized proteins (incomplete coverage).
  • Controls/quality steps: technical and biological replicates; spike-in standards for absolute quantification; false-discovery-rate (FDR) control in database matching (decoy database) to avoid mis-identifications; and validation by Western blot or targeted MS (e.g., SRM/PRM) of key hits.
  • Common mistakes: equating protein abundance with activity; over-trusting single-peptide identifications; ignoring FDR; incomplete digestion affecting quantitation; and assuming mRNA and protein levels will agree.

Common confusions

  • "Proteomics reads protein sequences like DNA sequencing" — MS identifies proteins by matching peptide masses/fragments to a database, not by base-by-base sequencing.
  • "mRNA levels predict protein levels" — They correlate poorly in many cases; translation, turnover, and PTMs intervene.
  • "The proteome is fixed by the genome" — One gene can yield many protein forms via splicing and PTMs; the proteome is dynamic and context-dependent.
  • "Abundance = activity" — A protein can be present but inactive (needs PTM, binding partner, or correct localization).
  • "MS sees every protein" — Low-abundance, hydrophobic, or poorly ionizing proteins are often missed (incomplete coverage).

Quick review

  • Extract proteins → trypsin digest → LC → MS1 (m/z) → MS2 (fragments) → database match → identify/quantify/PTMs.
  • Measures protein identity, abundance, and modifications — the functional layer.
  • Protein ≠ mRNA, and abundance ≠ activity; validate key hits with Western/targeted MS.
Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

Proteomics is taking inventory of the actual workers in a factory, not the instruction manuals. You can't guess who's working by counting manuals (mRNA) — some manuals are read a lot, some not. So you chop every worker into standard pieces, weigh each piece on a super-accurate scale (mass spectrometer), break them again, and match the pieces to a catalog to name each worker and count how many there are. (The analogy's limit: the inventory shows who is present, not who is actually working — an abundant protein can still be switched off.)

Key takeaways

  • ### High-Yield Facts
  • Proteomics = large-scale study of the proteome (proteins), measured by mass spectrometry.
  • Bottom-up workflow: digest (trypsin) → LC → MS1 (m/z) → MS2 (fragment) → database match.
  • Trypsin cuts after lysine (K) and arginine (R).
  • MS identifies by mass-to-charge ratio (m/z) + fragment spectra.
  • Detects PTMs (phosphorylation, etc.) via characteristic mass shifts.
  • Protein abundance ≠ activity; mRNA ≠ protein (translation/turnover/PTM).

Keep learning

Ready to build on this? Continue to the next lesson.

Study tools & related lessonsYou’ll learn to · Related

You’ll learn to

  • Define the proteome and explain why it cannot be read directly from the genome.
  • Describe how mass spectrometry identifies and quantifies proteins (bottom-up, MS/MS).
  • Explain how peptides are matched to a database to name proteins, and how PTMs are detected.
  • Distinguish proteomics from transcriptomics and from targeted methods like Western blotting.
  • Identify what proteomics can and cannot prove about protein function.

Sources & references

  1. NCI, "proteomics" (Dictionary of Genetics Terms). https://www.cancer.gov/publications/dictionaries/genetics-dictionary/def/proteomics
  2. NCI, "mass spectrometry" (Dictionary of Genetics Terms). https://www.cancer.gov/publications/dictionaries/genetics-dictionary/def/mass-spectrometry
  3. OpenStax, *Biology 2e*, "Genomics and Proteomics." https://openstax.org/books/biology-2e/pages/17-5-genomics-and-proteomics
  4. MedlinePlus, "What are proteins and what do they do?" https://medlineplus.gov/genetics/understanding/howgeneswork/protein/
  5. NCI, "biomarker" (Dictionary of Genetics Terms). https://www.cancer.gov/publications/dictionaries/genetics-dictionary/def/biomarker

This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.