Biology for AP Courses · Biotechnology and Genomics
Genomics and Proteomics
On this page 9 sections
In 30 seconds
Genomics Study of whole genomes — structure, function, evolution Full entry → is the study of entire genomes — their structure, function, and evolution. Proteomics Study of the complete set of proteins in a cell/organism Full entry → is the study of the Proteome The full set of proteins present at a given time Full entry →: the complete set of proteins produced by a cell, tissue, or organism at a given moment. The genome is like a fixed parts list; the proteome is what is actually being built and used right now. Because one gene can produce many different proteins, the proteome is larger and far more dynamic than the genome — which is exactly why knowing a DNA sequence alone cannot tell you what a cell is doing.
Why this matters
- Drug targets are proteins: most medicines act on proteins, so cataloging the proteins of a pathogen, tumor, or tissue suggests where drugs could act.
- Biomarkers: proteins whose abundance changes with disease can flag illness early — but candidate biomarkers must be validated in large studies (educational framing).
- Cell state, not just potential: the genome says what a cell could do; the proteome says what it is doing.
- Gene regulation: comparing genomes, transcriptomes, and proteomes reveals where regulation happens — at DNA, RNA, or protein level.
- AP® exam: central dogma plus Alternative splicing Joining different exon combinations to make multiple mRNAs from one gene Full entry →, post-translational modifications, and the "-omics" vocabulary are commonly tested.
The college version
Core Concepts
Structural genomics: the genome as blueprint
Structural genomics maps and sequences whole genomes, locates genes, and compares genomes across species (Comparative genomics Comparing genomes across species Full entry →). The logic is simple: genes that are conserved between distantly related species — bacteria, yeast, worms, flies, mice, humans — almost certainly do something important, and their function can often be studied in a convenient Model organism A species studied because it is convenient and informative (yeast, fly, worm, mouse — commonly taught) Full entry →. Commonly taught model organisms include E. coli, yeast, the nematode C. elegans, the fruit fly, zebrafish, and the mouse; their genes frequently have human counterparts whose jobs are easier to study in the lab.
Functional genomics: what genes actually do
Knowing a gene's sequence is not knowing its function. Functional genomics investigates activity. Transcriptomics measures which genes are transcribed into mRNA, using RNA-seq or DNA microarrays (chips with thousands of probes that detect specific mRNAs). Gene-disruption experiments — knockouts, or RNA interference to silence a gene — reveal what happens when a gene is missing. Large projects such as ENCODE have cataloged functional elements across the human genome, showing that much of the "noncoding" majority is active regulatory DNA (commonly taught finding).
From genome to transcriptome to proteome
Information flows DNA → RNA → protein, and each step adds complexity:
- The Transcriptome The complete set of mRNA molecules in a cell Full entry → is the complete set of mRNA molecules in a cell — which genes are switched on and how strongly.
- The proteome is the complete set of proteins — what was actually translated, modified, and allowed to persist.
Crucially, mRNA levels and protein levels are not perfectly correlated: an mRNA may be degraded before translation, or a protein may be modified or destroyed after synthesis. Regulation happens at every level, which is why "gene is transcribed" does not equal "protein is active."
Why the proteome is bigger than the genome
Humans have roughly 20,000–25,000 protein-coding genes (commonly taught estimate), yet the proteome is far larger, because one gene can generate many distinct proteins:
- Alternative splicing: during mRNA processing, different combinations of exons can be joined, so a single gene produces multiple mRNA variants and therefore multiple protein isoforms.
- Post-translational modifications (PTMs): after synthesis, proteins are chemically altered — commonly taught examples include phosphorylation, glycosylation, and acetylation — changing their activity, location, or lifetime.
- Interactions and complexes: proteins work in complexes, adding functional diversity beyond individual molecules.
The proteome also changes with cell type, developmental stage, and environment. It is a snapshot, not a constant — which is what makes it informative.
How proteomics is done
- Two-dimensional (2D) gel electrophoresis separates proteins in two passes — first by isoelectric point (charge), then by molecular mass — producing a field of spots; spots present in one sample but not another are candidates for further study.
- Mass spectrometry Identifies proteins by measuring ion mass-to-charge ratios Full entry → is the workhorse: it ionizes proteins or digested peptides and measures their mass-to-charge ratios, identifying proteins by mass and sequence with high sensitivity.
- Protein microarrays expose thousands of proteins or antibodies on a chip to detect expression or binding in parallel.
- Interaction screens such as the Yeast two-hybrid Technique that detects physical interactions between proteins Full entry → method reveal which proteins physically interact (commonly taught technique).
- Structure databases such as the Protein Data Bank (commonly referenced) store experimentally determined 3D structures of proteins.
Applications
Proteomics compares healthy and diseased samples to find Biomarker A measurable molecule whose change signals a condition Full entry → candidates — proteins whose abundance signals a condition — and catalogs the proteins of pathogens and tumors to propose drug targets. It also illuminates signaling: cascades of phosphorylation turn enzyme activity on and off, controlling everything from cell division to immune responses (commonly taught). Every candidate biomarker or target still requires rigorous validation before clinical or industrial use.
Common Confusions
| Do not confuse | With | Difference |
|---|---|---|
| Genomics | Genetics | Genome-wide study vs study of single genes |
| Genome | Proteome | The fixed parts list vs the proteins working right now |
| Transcriptome | Proteome | mRNA vs protein; degradation and regulation decouple the two |
| Alternative splicing | Mutation | Normal processing that generates multiple variants vs a change in the DNA sequence itself |
| mRNA level | Protein activity | A protein can be inactive despite being abundant (blocked by modification or inhibitor) |
| Structural genomics | Functional genomics | Where/what the genes are vs what the genes do |
| Proteomics | Individual protein structure study | Full-set analysis vs the 3D structure of one protein |
| Protein modification | Protein degradation | Altering a protein's chemistry vs destroying it |

Eli explains
The same idea, in plain words
Explain it like I’m 10
The genome is a cookbook with every recipe your body could make; the proteome is the meal actually on the table right now. The same recipe can become different dishes — add extra salt (a modification), skip an ingredient (a splice variant) — so the cookbook may be short, but the possible dishes are endless. Genomics reads the cookbook; proteomics looks at the table.
Worked example
The same genome, two very different proteomes. A surgeon removes a tumor and healthy tissue from the same person. The two samples' genomes are essentially identical, but their proteomes are not. Researchers run both samples on 2D gels: dozens of spots appear in the tumor lane only. Mass spectrometry identifies one of those spots as a protein 10 times more abundant in the tumor than in healthy tissue. Before anything is concluded, the candidate is tested in hundreds of additional patients — is it consistently elevated in this cancer? Does it predict outcome? If it survives validation, it may become a biomarker for detection or a drug target for therapy.
The lesson: the genome said what was possible; the proteome said what was happening — and in this case, the difference was the entire point. (Educational illustration of the research pipeline.)
Key takeaways
- Genomics = whole genomes; proteomics = the complete protein set (the proteome).
- One gene → many proteins: alternative splicing + post-translational modifications (phosphorylation, glycosylation, acetylation — commonly taught).
- The proteome is dynamic (changes with cell type, time, conditions); the genome is essentially fixed.
- Transcriptome ≠ proteome: mRNA and protein levels do not always match because regulation and degradation occur after transcription.
- Methods: 2D gels + mass spectrometry; protein microarrays; yeast two-hybrid for interactions.
- Model organisms reveal gene function because important genes are conserved across species.
- ENCODE: much of the "noncoding" genome is functional regulatory DNA (commonly taught).
- Proteins are the usual drug targets — proteomics powers drug discovery and biomarker research.
- Comparative genomics: conserved genes across species = important, ancient functions.
Check yourself
6 review questions from the chapter. Try each one, then open the answer.
Why is the proteome larger than the genome?
Show answer
Alternative splicing produces multiple mRNA variants (and protein isoforms) from one gene, and post-translational modifications further diversify proteins — so ~20,000–25,000 genes can encode vastly more distinct proteins.
What does alternative splicing accomplish?
Show answer
It joins different combinations of exons during mRNA processing, letting a single gene encode several different proteins.
Name two post-translational modifications (commonly taught examples).
Show answer
Phosphorylation and glycosylation (acetylation is another commonly taught example).
Why might a gene be transcribed yet produce little functional protein?
Show answer
The mRNA may be degraded before translation, translation may be blocked, or the protein may be modified or destroyed after synthesis — regulation happens at every step, so transcription alone guarantees nothing.
By what two properties does 2D gel electrophoresis Separates proteins by charge, then by mass Full entry → separate proteins?
Show answer
First by isoelectric point (charge), then by molecular mass — yielding a two-dimensional field of spots.
Why do drug-discovery programs study proteomes?
Show answer
Because most drugs act on proteins: cataloging the proteins a pathogen or tumor produces suggests which molecules to attack, and comparing proteomes finds disease biomarkers.
Study tools & related lessonsKey vocabulary · Related
Key vocabulary
- Genomics
- Study of whole genomes — structure, function, evolution
- Proteomics
- Study of the complete set of proteins in a cell/organism
- Proteome
- The full set of proteins present at a given time
- Transcriptome
- The complete set of mRNA molecules in a cell
- Alternative splicing
- Joining different exon combinations to make multiple mRNAs from one gene
- Post-translational modification
- Chemical change to a protein after synthesis (phosphorylation, glycosylation, acetylation — commonly taught)
- 2D gel electrophoresis
- Separates proteins by charge, then by mass
- Mass spectrometry
- Identifies proteins by measuring ion mass-to-charge ratios
- Protein microarray
- Chip carrying thousands of proteins/antibodies for parallel detection
- Yeast two-hybrid
- Technique that detects physical interactions between proteins
- Biomarker
- A measurable molecule whose change signals a condition
- Model organism
- A species studied because it is convenient and informative (yeast, fly, worm, mouse — commonly taught)
- Comparative genomics
- Comparing genomes across species
Sources & references
This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.

