Biology 1 · Genetics & Inheritance

Gene Expression: From DNA to Protein

33 min read
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 4 sections
  1. The college version
  2. Eli explains
  3. Key takeaway
  4. Study tools

The college version

Core Explanation

Gene expression is the process by which information encoded in a gene — a specific sequence of DNA — is used to direct the synthesis of a functional gene product, most commonly a protein. In all organisms, gene expression proceeds through two major stages: (DNA → RNA) and (RNA → protein). This flow of genetic information is the central dogma of molecular biology, first articulated by Francis Crick in 1958.

DNA  ──transcription──▶  RNA  ──translation──▶  Protein

Exceptions exist — some viruses use reverse transcription (RNA → DNA), and many RNA molecules (tRNA, rRNA, snRNA, microRNA) are functional without being translated — but the central dogma captures the fundamental information pathway for all cellular life.


Part I: Transcription — DNA-Directed RNA Synthesis

Transcription is the synthesis of an RNA molecule from a DNA template. The enzyme catalyzes this reaction. Only one of the two DNA strands — the (also called the antisense or noncoding strand) — serves as the template for RNA synthesis.

Key Players

ComponentRole
RNA polymeraseEnzyme that synthesizes RNA in the 5′ → 3′ direction by adding ribonucleoside triphosphates (NTPs) complementary to the DNA template. Does NOT require a primer.
Template strandThe DNA strand that RNA polymerase reads in the 3′ → 5′ direction. The synthesized RNA is complementary to this strand.
Coding strandThe non-template DNA strand. It has the same sequence as the RNA transcript (with T replaced by U). NOT used by RNA polymerase.
PromoterA specific DNA sequence upstream of the gene where RNA polymerase binds to initiate transcription. Defines the transcription start site and which DNA strand will be the template.
TerminatorA DNA sequence that signals RNA polymerase to stop transcription and release the transcript.

Critical distinction: The template strand is read 3′ → 5′ so the growing RNA chain is synthesized 5′ → 3′. The RNA transcript is complementary and antiparallel to the template strand. It is identical in sequence to the (with U replacing T):

Coding strand:     5′ — ATGCCGTTAGAC — 3′
Template strand:   3′ — TACGGCAATCTG — 5′
RNA transcript:    5′ — AUGCCGUUAGAC — 3′

Stages of Transcription

1. Initiation

In prokaryotes, RNA polymerase holoenzyme (core enzyme + ) recognizes and binds to the . The promoter contains conserved sequences at approximately −10 (TATAAT, the Pribnow box) and −35 (TTGACA) relative to the transcription start site (+1). The sigma factor confers promoter specificity; once transcription begins, sigma dissociates.

In eukaryotes, initiation is far more complex. RNA polymerase II (which transcribes protein-coding genes) requires a suite of general transcription factors (TFIID, TFIIB, TFIIF, TFIIE, TFIIH, etc.) to assemble at the promoter. TFIID contains the TATA-binding protein (TBP) that recognizes the TATA box (consensus TATAAAA, located about −25 to −30). This assembly forms the preinitiation complex. TFIIH has helicase activity that unwinds DNA and kinase activity that phosphorylates RNA polymerase II, triggering promoter escape.

Key point: RNA polymerase unwinds approximately 10–17 base pairs of DNA, forming a . The first few NTPs are linked together without primer requirement (unlike DNA polymerase).

2. Elongation

RNA polymerase moves along the template strand 3′ → 5′, adding NTPs one at a time to the 3′ end of the growing RNA chain. Incoming NTPs pair with complementary bases on the template (A–U, G–C, T–A, C–G). A phosphodiester bond forms between the 3′-OH of the growing chain and the 5′-phosphate of the incoming NTP, releasing pyrophosphate (PPi).

The transcription bubble moves with the polymerase. DNA unwinds ahead and rewinds behind the bubble. The RNA transcript peels away from the template, and only about 8–9 nucleotides remain base-paired with DNA at any moment.

In eukaryotes, RNA polymerase II begins transcribing beyond the eventual 3′ end of the mRNA. Additional processing signals are encountered downstream.

3. Termination

Prokaryotic termination occurs by two mechanisms:

  • Rho-independent (intrinsic) termination: The RNA transcript forms a GC-rich hairpin loop followed by a string of U residues. The hairpin causes RNA polymerase to pause, and the weak A–U base pairing (only 2 hydrogen bonds between A and U in the RNA–DNA hybrid, compared to 3 for G–C) allows the transcript to dissociate.
  • Rho-dependent termination: The rho (ρ) protein, an ATP-dependent helicase, binds to a rut (rho utilization) site on the nascent RNA, translocates along the transcript, and unwinds the RNA–DNA hybrid when it catches up to the paused polymerase.

Eukaryotic termination: For protein-coding genes transcribed by RNA polymerase II, termination is coupled to RNA processing. After RNA polymerase transcribes past the polyadenylation signal (AAUAAA in the RNA), the transcript is cleaved ~10–35 nucleotides downstream. RNA polymerase II continues transcribing for hundreds of nucleotides before eventually dissociating; the trailing RNA is degraded. The cleaved pre-mRNA then undergoes processing.


Part II: Eukaryotic RNA Processing

In prokaryotes, transcription and translation are coupled: ribosomes begin translating mRNA while it is still being transcribed. In eukaryotes, the primary transcript (pre-mRNA) must undergo extensive processing in the nucleus before it can be exported to the cytoplasm for translation.

The Three Major Processing Events

1. 5′ Capping

Shortly after transcription begins (when the RNA is ~20–30 nucleotides long), a 7-methylguanosine cap is added to the 5′ end of the pre-mRNA via a 5′–5′ triphosphate linkage. The capping enzyme adds this modified guanine nucleotide in an unusual orientation.

Functions of the 5′ cap:

  • Protects the mRNA from degradation by 5′ exonucleases
  • Facilitates transport from the nucleus to the cytoplasm
  • Recognized by the translation initiation complex (cap-binding proteins) for ribosome recruitment
2. 3′ Polyadenylation (Poly-A Tail)

After cleavage at the polyadenylation signal, the enzyme poly-A polymerase adds approximately 50–250 adenine nucleotides (the ) to the new 3′ end. This addition does NOT use a DNA template — it is a post-transcriptional modification.

Functions of the poly-A tail:

  • Protects the mRNA from 3′ exonuclease degradation
  • Facilitates nuclear export
  • Enhances translation initiation (poly-A binding proteins interact with cap-binding proteins, forming a closed-loop structure)
  • The tail shortens over time in the cytoplasm; when it becomes too short, the mRNA is degraded — this is a mechanism for controlling mRNA lifespan
3. Splicing: Intron Removal

Most eukaryotic genes are interrupted by introns (intervening sequences) that do not code for protein. The coding sequences — exons — are expressed. Introns are removed and exons are joined together in a process called splicing, which occurs in the nucleus.

The :

The spliceosome is a large RNA–protein complex that catalyzes splicing. It is composed of five small nuclear ribonucleoproteins (snRNPs, pronounced "snurps"): U1, U2, U4, U5, and U6. Each snRNP contains a small nuclear RNA (snRNA) and associated proteins. The snRNAs recognize the intron boundaries through complementary base pairing.

Splicing mechanism:

  1. U1 snRNP binds the 5′ splice site (GU dinucleotide at the intron's 5′ end).
  2. U2 snRNP binds the branch point adenine within the intron.
  3. U4/U6 and U5 snRNPs join, forming the active spliceosome.
  4. The 5′ end of the intron is cleaved and covalently linked to the branch point adenine, forming a lariat structure.
  5. The 3′ splice site (AG dinucleotide) is cleaved, and the two exons are ligated together.
  6. The intron lariat is released and degraded.

The GU–AG rule: nearly all introns begin with GU and end with AG. These conserved sequences are essential for splice site recognition.

What makes splicing remarkable: The spliceosome is a ribozyme — the catalytic activity comes from the snRNAs, not the proteins. This is a key piece of evidence for the RNA world hypothesis.

Alternative Splicing

allows a single gene to produce multiple different mRNA transcripts (and thus multiple protein isoforms) by joining different combinations of exons. This is a primary explanation for why the ~20,000 protein-coding genes in the human genome can generate an estimated 100,000+ distinct proteins.

Types of alternative splicing include:

  • Exon skipping (cassette exon): an exon is either included or excluded.
  • Alternative 5′ or 3′ splice sites: the splice site boundary shifts, altering exon length.
  • Intron retention: an intron is retained in the mature mRNA (common in plants, rarer in animals).
  • Mutually exclusive exons: one of two exons is included, but not both.

Alternative splicing is regulated by RNA-binding proteins (splicing factors) that enhance (splicing enhancers) or repress (splicing silencers) splice site recognition. This regulation can be tissue-specific, developmental-stage-specific, or responsive to cellular signals.

Example: The Dscam gene in Drosophila has 38,016 possible alternative splice isoforms — more than the total number of genes in the fly genome. In humans, the troponin T gene produces different isoforms in cardiac and skeletal muscle through alternative splicing.

Pre-mRNA Processing Summary

Pre-mRNA:  5' ──[exon 1]──[==intron 1==]──[exon 2]──[==intron 2==]──[exon 3]── 3'

Processing:  Add 5' cap → Splice out introns → Cleave and add poly-A tail

Mature mRNA:  5'-cap ──[exon 1]──[exon 2]──[exon 3]──AAAAAAAAAAA-3'

Part III: Translation — RNA-Directed Protein Synthesis

Translation is the synthesis of a polypeptide chain using the information encoded in mRNA. It occurs on ribosomes in the cytoplasm (or on the rough ER). The language of nucleic acids (nucleotides) is translated into the language of proteins (amino acids).

The Genetic Code

The genetic code is the set of rules by which nucleotide triplets (codons) specify amino acids. Each codon consists of three consecutive nucleotides in the mRNA (read 5′ → 3′).

Fundamental properties of the genetic code:

PropertyDescription
Triplet codeThree nucleotides = one codon = one amino acid (or stop signal)
Non-overlappingNucleotides are read in consecutive, non-overlapping groups of three
CommalessNo punctuation between codons; the reading frame is maintained continuously
Degenerate (redundant)Most amino acids are specified by more than one codon (see below)
UnambiguousEach codon specifies only ONE amino acid (no codon codes for two different amino acids)
Nearly universalShared by almost all organisms (minor exceptions in mitochondria and some protists)
Start codonAUG — codes for methionine; also serves as the initiation signal
Stop codonsUAA, UAG, UGA — do not code for any amino acid; signal termination

Of the 64 possible codons: 61 code for amino acids, and 3 are stop codons.

The Genetic Code Table

          Second base
          U         C         A         G
     ┌─────────┬─────────┬─────────┬─────────┐
  U  │ UUU Phe │ UCU Ser │ UAU Tyr │ UGU Cys │ U
     │ UUC Phe │ UCC Ser │ UAC Tyr │ UGC Cys │ C
     │ UUA Leu │ UCA Ser │ UAA STOP│ UGA STOP│ A
     │ UUG Leu │ UCG Ser │ UAG STOP│ UGG Trp │ G
     ├─────────┼─────────┼─────────┼─────────┤
  C  │ CUU Leu │ CCU Pro │ CAU His │ CGU Arg │ U
F    │ CUC Leu │ CCC Pro │ CAC His │ CGC Arg │ C
i    │ CUA Leu │ CCA Pro │ CAA Gln │ CGA Arg │ A
r    │ CUG Leu │ CCG Pro │ CAG Gln │ CGG Arg │ G
s    ├─────────┼─────────┼─────────┼─────────┤
t  A │ AUU Ile │ ACU Thr │ AAU Asn │ AGU Ser │ U
     │ AUC Ile │ ACC Thr │ AAC Asn │ AGC Ser │ C
b    │ AUA Ile │ ACA Thr │ AAA Lys │ AGA Arg │ A
a    │ AUG Met │ ACG Thr │ AAG Lys │ AGG Arg │ G
s    ├─────────┼─────────┼─────────┼─────────┤
e  G │ GUU Val │ GCU Ala │ GAU Asp │ GGU Gly │ U
     │ GUC Val │ GCC Ala │ GAC Asp │ GGC Gly │ C
     │ GUA Val │ GCA Ala │ GAA Glu │ GGA Gly │ A
     │ GUG Val │ GCG Ala │ GAG Glu │ GGG Gly │ G
     └─────────┴─────────┴─────────┴─────────┘

Degeneracy and the Wobble Hypothesis

The code is degenerate: 18 of the 20 amino acids are encoded by more than one codon. Methionine and tryptophan are the only exceptions, each having a single codon (AUG and UGG, respectively).

Codons for the same amino acid tend to differ only in the third base position — often called the "wobble" position. The (proposed by Francis Crick in 1966) explains why fewer tRNAs are needed than there are codons: the base pairing between the third position of the codon (mRNA) and the first position of the anticodon (tRNA) is less stringent — nonstandard base pairing (e.g., G–U, inosine–U/C/A) is permitted.

This means a single tRNA can recognize multiple codons for the same amino acid. For example, there are six codons for leucine but far fewer leucine-specific tRNAs, because wobble pairing allows one tRNA anticodon to pair with multiple leucine codons.

Why degeneracy matters: It buffers against the effects of point mutations. A change in the third codon position often produces a silent mutation — the same amino acid is still incorporated.

Transfer RNA (tRNA): The Adaptor Molecule

tRNA is the molecular adaptor that translates the nucleic acid language of mRNA into the amino acid language of proteins. Each tRNA molecule has two critical functional regions:

  1. Anticodon: A three-nucleotide sequence that base-pairs (antiparallel) with a complementary mRNA codon.
  2. Acceptor stem (3′ end): The sequence CCA at the 3′ terminus, where the corresponding amino acid is covalently attached.

tRNA folds into a characteristic cloverleaf secondary structure (in 2D) and an L-shaped tertiary structure (in 3D). This precise shape is essential for fitting into the ribosome's A and P sites.

Aminoacyl-tRNA Synthetase

Each of the 20 amino acids has its own aminoacyl-tRNA synthetase — an enzyme that catalyzes the attachment of the correct amino acid to the correct tRNA in a two-step reaction:

Amino acid + ATP → Aminoacyl-AMP + PPi
Aminoacyl-AMP + tRNA → Aminoacyl-tRNA + AMP

The charged tRNA is now called an aminoacyl-tRNA. The high-energy bond between the amino acid and the tRNA provides the energy for peptide bond formation later.

Proofreading: Aminoacyl-tRNA synthetases have editing sites that hydrolyze incorrectly attached amino acids. This double-check mechanism is essential — an error rate of even 1 in 10,000 at this step would introduce errors into nearly every protein. The overall accuracy of aminoacyl-tRNA synthetases exceeds 99.99%.

The Ribosome: The Protein Synthesis Factory

Ribosomes are large ribonucleoprotein complexes composed of two subunits.

Prokaryotic RibosomeEukaryotic Ribosome
Overall70S80S
Large subunit50S (23S rRNA + 5S rRNA + ~31 proteins)60S (28S rRNA + 5.8S rRNA + 5S rRNA + ~49 proteins)
Small subunit30S (16S rRNA + ~21 proteins)40S (18S rRNA + ~33 proteins)

Three tRNA binding sites are formed at the interface of the two subunits:

SiteFull NameFunction
A siteAminoacyl siteAccepts the incoming aminoacyl-tRNA (except initiator tRNA)
P sitePeptidyl siteHolds the tRNA carrying the growing polypeptide chain
E siteExit siteHolds the deacylated tRNA (now without an amino acid) before it exits the ribosome

Critical: The peptidyl transferase activity that forms peptide bonds is catalyzed by the 23S rRNA (prokaryotes) or 28S rRNA (eukaryotes) in the large subunit — not by any ribosomal protein. The ribosome is a ribozyme. This discovery (Steitz, Yonath, Ramakrishnan — Nobel Prize in Chemistry, 2009) provided powerful evidence for the RNA world hypothesis.

Stages of Translation

1. Initiation

Prokaryotic initiation:

  1. The 30S small subunit binds to the Shine–Dalgarno sequence (AGGAGG) in the mRNA, located ~5–10 nucleotides upstream of the start codon (AUG). This sequence is complementary to the 3′ end of 16S rRNA, positioning the start codon in the P site.
  2. The initiator tRNA (carrying N-formylmethionine, fMet) binds to the AUG start codon in the P site with the help of initiation factors (IF1, IF2, IF3).
  3. The 50S large subunit joins, forming the 70S initiation complex. Initiation factors are released, and GTP is hydrolyzed.

Eukaryotic initiation (simplified):

  1. The 40S small subunit, bound to initiation factors and the initiator tRNA (carrying methionine, not formylated), scans the mRNA from the 5′ cap until it encounters an AUG in a favorable sequence context (Kozak sequence: GCCRCCAUGG, where R = purine).
  2. Upon AUG recognition, the 60S large subunit joins, forming the 80S initiation complex.
  3. Note: eukaryotic initiation is FAR more complex — over a dozen initiation factors are involved (eIFs).

Key difference: Prokaryotes use a Shine–Dalgarno sequence for ribosome binding; eukaryotes use cap-dependent scanning.

2. Elongation

Elongation is a cyclic three-step process (repeated for each amino acid added):

Step 1 — Codon recognition (tRNA binding): An aminoacyl-tRNA with an anticodon complementary to the codon in the A site enters the A site. This process requires elongation factor EF-Tu (prokaryotes) or eEF1α (eukaryotes) and GTP hydrolysis for proofreading.

Step 2 — Peptide bond formation (transpeptidation): The peptidyl transferase center (the rRNA of the large subunit) catalyzes the formation of a peptide bond between the amino group of the incoming amino acid (in the A site) and the carboxyl end of the growing chain (attached to the tRNA in the P site). The polypeptide chain is transferred from the tRNA in the P site to the tRNA in the A site. No additional energy input is needed — the energy comes from the high-energy aminoacyl-tRNA bond.

Step 3 — Translocation: The ribosome moves (translocates) one codon toward the 3′ end of the mRNA. This requires elongation factor EF-G (prokaryotes) or eEF2 (eukaryotes) and GTP hydrolysis. After translocation:

  • The deacylated tRNA (now uncharged) moves from the P site to the E site, then exits.
  • The peptidyl-tRNA (with the growing chain) moves from the A site to the P site.
  • The A site is now empty and aligned with the next codon, ready for the next aminoacyl-tRNA.

Energy cost per amino acid added: 2 GTP (one for tRNA binding, one for translocation) + 1 ATP equivalent (from the aminoacyl-tRNA charging reaction) = 3 high-energy phosphate bonds per peptide bond.

3. Termination

Termination occurs when a stop codon (UAA, UAG, or UGA) enters the A site:

  1. No tRNA recognizes stop codons. Instead, a protein release factor (RF1 or RF2 in prokaryotes; eRF1 in eukaryotes) binds the stop codon in the A site. Release factors have a shape that mimics tRNA.
  2. The release factor triggers the peptidyl transferase center to hydrolyze the bond linking the polypeptide chain to the tRNA in the P site (adding a water molecule instead of an amino acid).
  3. The polypeptide chain is released from the ribosome.
  4. The ribosomal subunits, mRNA, tRNA, and release factors dissociate (aided by GTP hydrolysis and ribosome recycling factors).

Polysomes (polyribosomes): Multiple ribosomes can simultaneously translate a single mRNA molecule, each at a different stage of translation. A typical eukaryotic mRNA might have 5–10 ribosomes translating it at once. This allows a single mRNA to produce many protein copies before it is degraded.

Translation Summary Table

StageKey Events
InitiationSmall subunit binds mRNA → initiator tRNA pairs with AUG at P site → large subunit joins
ElongationAminoacyl-tRNA enters A site → peptide bond forms (P → A) → translocation (A → P, P → E) → repeat
TerminationStop codon enters A site → release factor binds → polypeptide released → ribosome dissociates

Part IV: Gene Regulation — An Overview

Not all genes are expressed at all times. Cells regulate which genes are transcribed and translated, controlling both the timing and the amount of gene expression. This regulation is essential for development, differentiation, homeostasis, and response to environmental signals.

Prokaryotic Gene Regulation: Operons

Prokaryotes often organize functionally related genes into operons — clusters of genes under the control of a single promoter that are transcribed together as a single polycistronic mRNA.

The lac Operon: Inducible (Catabolic) System

The lac operon in E. coli contains three structural genes (lacZ, lacY, lacA) encoding enzymes for lactose metabolism: β-galactosidase (lacZ), lactose permease (lacY), and thiogalactoside transacetylase (lacA). These genes are transcribed as a single mRNA.

Regulatory components:

  • Promoter (P): Binding site for RNA polymerase.
  • Operator (O): DNA sequence between the promoter and structural genes; the binding site for the lac repressor.
  • lacI gene: Located upstream of the operon; constitutively expressed. It encodes the lac repressor protein.
  • CAP binding site: Upstream of the promoter; binding site for catabolite activator protein (CAP).

Regulation logic:

ConditionRepressorCAPTranscription
No lactose, glucose presentBound to operatorInactive (no cAMP)OFF
Lactose, glucose presentInactive (bound to allolactose)Inactive (no cAMP)Low (basal)
Lactose, no glucoseInactive (bound to allolactose)Active (cAMP bound, CAP binds)HIGH
No lactose, no glucoseBound to operatorActive (CAP bound)OFF (repressor blocks)

The dual control ensures efficiency: The operon is only fully ON when lactose is available AND glucose is absent. Glucose is the preferred energy source — if glucose is present, it is energetically wasteful to produce lactose-metabolizing enzymes.

How allolactose deactivates the repressor: A small amount of lactose is converted to allolactose (the inducer) by existing β-galactosidase. Allolactose binds to the lac repressor, causing a conformational change that prevents the repressor from binding the operator. This is an inducible system — the substrate (lactose) induces enzyme synthesis.

The trp Operon: Repressible (Anabolic) System

The trp operon encodes five enzymes for tryptophan biosynthesis.

Regulatory logic: Unlike the lac operon (which is induced by its substrate), the trp operon is repressed by its product. When tryptophan is abundant, the cell does not need to synthesize more.

Mechanism:

  • The trp repressor is produced in an inactive form by a separate regulatory gene (trpR).
  • When tryptophan levels are high, tryptophan acts as a corepressor — it binds to the trp repressor, activating it.
  • The active repressor–tryptophan complex binds to the operator, blocking transcription.
ConditionRepressorTranscription
Low tryptophanInactive (no corepressor bound)ON
High tryptophanActive (tryptophan corepressor bound)OFF

Attenuation: The trp operon also uses a second regulatory mechanism called attenuation, which couples transcription and translation. When tryptophan is abundant, ribosomes translate the leader peptide quickly, causing the formation of a transcription terminator hairpin in the mRNA — premature termination. Attenuation allows fine-tuning beyond simple on/off control.

Summary: Inducible vs Repressible Operons

Featurelac Operontrp Operon
Pathway typeCatabolic (breakdown)Anabolic (synthesis)
Regulation typeInducibleRepressible
Default stateOFFON
Effector moleculeAllolactose (inducer)Tryptophan (corepressor)
Effector effectInactivates repressor → ONActivates repressor → OFF
Responds toPresence of substrateAbundance of product

Exam tip: Inducible systems (lac, ara) control catabolic pathways — turn ON when substrate is present. Repressible systems (trp, his) control anabolic pathways — turn OFF when product is abundant. This makes energetic sense.


Eukaryotic Gene Regulation

Eukaryotic gene regulation is vastly more complex than prokaryotic regulation. Every step in the gene expression pathway can be controlled — from chromatin accessibility through transcription, RNA processing, mRNA export, translation, and protein degradation. Below is an overview of the major mechanisms, focused primarily on transcriptional regulation.

1. Chromatin Structure and Accessibility

In eukaryotes, DNA is packaged with histone proteins into chromatin. The default state is "packed up" and inaccessible. Gene activation requires chromatin remodeling to expose promoter regions.

Euchromatin: Loosely packed, transcriptionally active. \ Heterochromatin: Tightly packed, transcriptionally silent. (Includes constitutive heterochromatin — always silent, such as centromeres and telomeres — and facultative heterochromatin, which can be reactivated depending on the cell type or developmental stage.)

Chromatin remodeling complexes (e.g., SWI/SNF) use ATP hydrolysis to slide or eject nucleosomes, exposing DNA to transcription factors and RNA polymerase.

2. Histone Modification

The N-terminal tails of histone proteins protrude from the nucleosome core and are subject to various covalent modifications that affect chromatin structure and gene expression:

ModificationEffect on TranscriptionKey Enzymes
Histone acetylation (adding acetyl groups to lysine)Activates transcription — neutralizes positive charge on lysine, weakening histone–DNA interaction (opens chromatin)Histone acetyltransferases (HATs); removed by histone deacetylases (HDACs)
Histone methylation (adding methyl groups to lysine or arginine)Context-dependent: can activate OR repress depending on which residue is methylatedHistone methyltransferases (HMTs); removed by histone demethylases
Histone phosphorylationOften associated with chromatin condensation during mitosis and transcriptional activation at specific genesKinases; removed by phosphatases

The histone code hypothesis: Combinations of histone modifications create a "code" that is read by other proteins to determine the transcriptional state of a gene. For example, H3K4me3 (trimethylation of lysine 4 on histone H3) is strongly associated with active promoters, while H3K27me3 is associated with gene silencing.

3. DNA Methylation

DNA methylation — the addition of a methyl group to the 5′ position of cytosine, typically in CpG dinucleotides — is generally associated with transcriptional repression:

  • Methylated CpG sequences in promoter regions recruit methyl-CpG-binding proteins, which in turn recruit histone deacetylases (HDACs) and other repressive chromatin modifiers.
  • CpG islands (GC-rich regions near promoters) are usually unmethylated in active genes and methylated in silenced genes.

DNA methylation patterns are heritable through cell division (maintained by DNA methyltransferases like DNMT1), providing a mechanism for epigenetic inheritance — stable changes in gene expression that do NOT involve changes in the DNA sequence itself.

Examples: X-chromosome inactivation in female mammals and genomic imprinting both involve DNA methylation. Aberrant DNA methylation is also a hallmark of many cancers — tumor suppressor genes are often hypermethylated (silenced), while oncogenes may be hypomethylated (activated).

4. Enhancers and Silencers

Enhancers are DNA sequences (typically 50–1,500 bp) that can be located thousands of base pairs away from the promoter — upstream, downstream, within introns, or even on a different chromosome (in trans). They bind activator proteins (specific transcription factors) and increase transcription of associated genes.

Silencers are analogous DNA sequences that bind repressor proteins and decrease transcription.

How do enhancers work over long distances? DNA looping brings enhancer-bound activator proteins into physical contact with the promoter-bound transcription machinery. The intervening DNA loops out, and mediator complexes and cohesin help bridge and stabilize the interaction.

Key properties of enhancers:

  • They work in an orientation-independent manner (can be inverted and still function).
  • They can act at great distances (some enhancers are >1 Mb from their target gene).
  • They are cell-type-specific — a given enhancer is active only in cell types where the appropriate activator proteins are present.
  • Insulators (boundary elements) block enhancer–promoter interactions, defining regulatory domains and preventing inappropriate gene activation.
5. Transcription Factors

Transcription factors are proteins that bind specific DNA sequences and regulate transcription.

General (basal) transcription factors: Required for transcription by RNA polymerase II at ALL promoters (e.g., TFIID, TFIIB, TFIIH). They are part of the basic transcriptional machinery — necessary but not sufficient for high-level, regulated expression.

Specific transcription factors (regulatory): Bind to enhancers, silencers, or promoter-proximal elements and regulate transcription of SPECIFIC genes. They contain at least two functional domains:

  • DNA-binding domain: Recognizes a specific DNA sequence (e.g., helix-turn-helix, zinc finger, leucine zipper, helix-loop-helix motifs).
  • Activation domain (or repression domain): Interacts with other proteins (mediator complex, chromatin modifiers, general transcription factors) to stimulate or repress transcription.
6. Combinatorial Control

A single gene is typically regulated by many different transcription factors acting in combination. Conversely, a single transcription factor can regulate many different genes. This combinatorial control allows a limited number of transcription factors (humans have ~1,600) to regulate >20,000 genes in precise, tissue-specific, and signal-responsive patterns.

How combinatorial control works:

  • Multiple activator proteins must bind simultaneously for a gene to be transcribed (AND logic).
  • Different combinations of transcription factors activate different target genes in different cell types.
  • The same transcription factor can activate one gene and repress another, depending on its binding partners at each locus.

Example: The human β-globin gene is regulated by at least five different transcription factors binding to its enhancer and promoter. The specific combination of factors present in erythroid precursor cells (but not in other cell types) drives high-level expression — even though some of those factors are present in other tissues, the full complement is only present in red blood cell precursors.


Prokaryotic vs Eukaryotic Gene Regulation at a Glance

FeatureProkaryotesEukaryotes
OrganizationOperons (polycistronic mRNAs)Individual genes (mostly monocistronic)
ChromatinNo (naked DNA, though nucleoid-associated proteins exist)Yes — chromatin structure is a major regulatory layer
Primary control levelTranscription initiation (operators, repressors, activators)Transcription initiation (chromatin, enhancers, TFs) but also RNA processing, export, mRNA stability, translation
Regulatory proteinsRepressors and activators (simple)Hundreds of transcription factors + coactivators, corepressors, chromatin modifiers
DNA loopingCAP bends DNA; AraC loops DNAExtensive long-range enhancer–promoter looping
Post-transcriptionalMinimal (mRNA is short-lived)Extensive (alternative splicing, RNA editing, RNAi, mRNA localization, differential stability)

Biological / Medical Relevance

  • Antibiotics targeting transcription and translation: Rifampin binds bacterial RNA polymerase (treats tuberculosis). Tetracycline blocks tRNA binding to the A site; chloramphenicol inhibits peptidyl transferase; erythromycin blocks translocation. These selectively target prokaryotic ribosomes (70S), minimizing host toxicity. Streptomycin causes misreading of the genetic code.
  • Splicing defects and disease: ~15% of human genetic diseases are caused by mutations that disrupt splicing. Spinal muscular atrophy (SMA) results from mutations affecting SMN2 splicing. The drug nusinersen (Spinraza) is an antisense oligonucleotide that corrects SMN2 splicing.
  • Cancer and gene regulation: Aberrant DNA methylation (hypermethylation of tumor suppressors, hypomethylation of oncogenes), histone modification defects, and mutations in chromatin remodeling complexes (e.g., SWI/SNF) are hallmarks of cancer.
  • Epigenetics and development: Genomic imprinting disorders (Prader-Willi, Angelman syndromes) result from defects in DNA methylation at imprinted loci. X-inactivation in female mammals is an epigenetic phenomenon governed by the lncRNA XIST.
  • CRISPR-Cas9 gene editing: Engineered transcription factors (CRISPR activation, CRISPR interference) can be targeted to specific promoters or enhancers to activate or repress endogenous genes — directly leveraging the gene regulation principles described here.
  • COVID-19 mRNA vaccines: The Pfizer-BioNTech and Moderna vaccines deliver synthetic mRNA with a 5′ cap and poly-A tail, optimized codons, and modified nucleosides — a direct clinical application of the gene expression machinery.

Common Misconceptions and Exam Traps

  • Exam trap: The template strand is read 3′ → 5′, and RNA is synthesized 5′ → 3′. Students routinely get this backward. Think: polymerases ALWAYS synthesize 5′ → 3′.
  • Exam trap: "The coding strand is the template." WRONG. The coding strand has the SAME sequence as RNA (T → U); the template strand is the one that RNA polymerase reads.
  • Misconception: "RNA polymerase needs a primer." RNA polymerase does NOT require a primer, unlike DNA polymerase. This is a classic exam distinction.
  • Exam trap: "Ribosomal proteins catalyze peptide bond formation." WRONG. The ribosome is a ribozyme — the 23S/28S rRNA catalyzes peptide bond formation. This is a Nobel Prize-level concept that exams love.
  • Misconception: "The lac operon is always ON when lactose is present." False. With glucose present, cAMP is low, CAP is inactive, and transcription is only at a basal (low) level. The lac operon requires BOTH lactose present AND glucose absent for full activation.
  • Exam trap: Confusing lac (inducible → substrate turns it ON) with trp (repressible → product turns it OFF). An easy mnemonic: inducible = catabolic (breakdown); repressible = anabolic (synthesis). Makes energetic sense.
  • Misconception: "Degeneracy means the genetic code is ambiguous." WRONG. The code is degenerate (multiple codons per amino acid) but UNAMBIGUOUS (each codon codes for only ONE amino acid). These are distinct properties — exams test this distinction.
  • Exam trap: "The stop codon is recognized by a special tRNA." WRONG. Stop codons are recognized by protein release factors (RF1/RF2/eRF1), NOT by tRNAs.
  • Misconception: "Alternative splicing means the same protein is made in different ways." Wrong direction — alternative splicing means different mRNA transcripts (and thus DIFFERENT proteins) are made from the SAME gene.
  • Misconception: "DNA methylation always silences genes." While generally repressive, DNA methylation in gene bodies (as opposed to promoters) can sometimes be associated with active transcription. The promoter context matters.

Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

Your DNA is a giant recipe book, but it stays locked in the nucleus (the "library" of the cell). When the cell needs to make a protein, it can't take the original recipe out — that's too risky. So it makes a copy of just the page it needs. That's transcription — copying a gene from DNA into messenger RNA. Before the copy leaves the nucleus, the cell edits it: it adds a protective cap to the front, a tail of extra letters to the back, and cuts out the "nonsense" parts (introns), gluing the useful parts (exons) together. The finished message heads out to the cytoplasm, where ribosomes — tiny machines made of RNA and protein — read the message three letters at a time. Each three-letter word (codon) means one amino acid. Special carriers called tRNA bring the right amino acid for each codon, and the ribosome links them into a chain. That chain folds up into a working protein. The cell also has ways to decide WHICH recipes to copy, and how many times — that's gene regulation. Prokaryotic cells use simple switches (operons); eukaryotic cells have a whole control room full of regulators, chemical tags on DNA and histones, and remote on/off switches called enhancers.


Key takeaways

  • The central dogma: DNA → RNA → Protein. Transcription produces RNA; translation produces protein.
  • RNA polymerase reads the template strand 3′ → 5′; RNA is synthesized 5′ → 3′. No primer required.
  • Promoters define the transcription start site and DNA strand to be transcribed. Eukaryotes use TATA box + general transcription factors; prokaryotes use −10 and −35 consensus sequences.
  • Eukaryotic pre-mRNA processing: 5′ cap (protection + ribosome recruitment), splicing (intron removal by spliceosome), poly-A tail (protection + export).
  • The spliceosome is a ribozyme — catalytic activity comes from snRNA, not protein.
  • Alternative splicing: one gene → multiple protein isoforms. Explains human proteome complexity.
  • Genetic code: triplet, non-overlapping, degenerate, unambiguous, nearly universal.
  • Start codon = AUG (Met); stop codons = UAA, UAG, UGA.
  • Wobble: relaxed third-position base pairing allows fewer tRNAs to cover all codons.
  • tRNA is the adaptor; aminoacyl-tRNA synthetase charges tRNA with correct amino acid (proofreads).
  • Ribosome A site (incoming tRNA), P site (peptidyl-tRNA), E site (exit). Peptide bonds are catalyzed by rRNA (ribozyme).
  • Translation: initiation (small subunit + initiator tRNA at AUG → large subunit joins), elongation (codon recognition → peptide bond → translocation × N), termination (release factor at stop codon).
  • Energy cost: ~3 high-energy phosphate bonds per peptide bond formed.
  • lac operon: inducible, catabolic. Default OFF; induced by allolactose (inactivates repressor). Also subject to positive control by CAP (glucose sensing).
  • trp operon: repressible, anabolic. Default ON; repressed by tryptophan (corepressor activates repressor).
  • Eukaryotic regulation is multilayered: chromatin remodeling, histone modifications (acetylation = ON), DNA methylation (generally OFF), enhancers/silencers, combinatorial TF control.
  • Enhancers act at a distance, orientation-independent, and are cell-type-specific.
  • ---
  • Transcription: RNA polymerase binds promoter → reads template strand 3′ → 5′ → synthesizes RNA 5′ → 3′ (initiation → elongation → termination). No primer needed.
  • Eukaryotic RNA processing: 5′ cap (protection) + intron splicing (spliceosome — a ribozyme) + poly-A tail (protection + export).
  • Genetic code: Triplet, degenerate, unambiguous, nearly universal. AUG = Met (start); UAA/UAG/UGA = stop.
  • tRNA: Anticodon pairs with codon; aminoacyl-tRNA synthetase charges tRNA (proofreads). Wobble at 3rd position.
  • Ribosome: A site (incoming), P site (peptide), E site (exit). Peptidyl transferase = rRNA (ribozyme).
  • Translation: Initiation (small subunit + initiator tRNA → large subunit) → elongation (3-step cycle: binding, peptide bond, translocation) → termination (release factor at stop codon).
  • lac operon: Inducible (substrate → ON). trp operon: Repressible (product → OFF).
  • Eukaryotic regulation: Chromatin (euchromatin/heterochromatin), histone acetylation (ON), DNA methylation (OFF), enhancers/silencers, combinatorial TF control.
  • ---
  • A gene has the following template strand sequence (partial): 3′ — TACGATC — 5′. What is the sequence of the RNA transcript produced from this region?
  • Why does the lac operon NOT produce high levels of β-galactosidase when both glucose and lactose are present?
  • A mutation changes a codon from GAA to GAG. Both code for glutamic acid. What type of mutation is this, and why is it possible?
  • How many high-energy phosphate bonds are consumed for each amino acid added to a growing polypeptide chain during elongation (including the energy cost of tRNA charging)?
  • Contrast the default state and regulatory logic of the lac operon with the trp operon. Why does it make energetic sense that catabolic operons are inducible and anabolic operons are repressible?
  • ---
  • Template strand (read 3′ → 5′): 3′ — TACGATC — 5′. RNA is synthesized 5′ → 3′, complementary to the template (A → U, T → A, C → G, G → C). The RNA transcript is: 5′ — AUGCUAG — 3′. (This is the same sequence as the coding strand, with T replaced by U.)
  • The lac operon requires BOTH the absence of the repressor AND the presence of active CAP for high-level transcription. With glucose present, intracellular cAMP levels are low, so CAP remains inactive and cannot bind its site to recruit RNA polymerase. The lac repressor is inactivated by allolactose, so there is low (basal) transcription, but without CAP-mediated positive control, transcription is NOT fully activated. The cell conserves energy by preferentially using glucose.
  • This is a silent mutation — a nucleotide change that does not alter the amino acid sequence of the protein. It is possible because of the degeneracy (redundancy) of the genetic code: both GAA and GAG specify glutamic acid. The change occurred at the third (wobble) codon position, where base-pairing is less stringent.
  • Three high-energy phosphate bonds per amino acid: one from ATP during tRNA charging (aminoacyl-tRNA synthetase: amino acid + ATP → aminoacyl-AMP + PPi; this is equivalent to 2 high-energy bonds, but the conversion of ATP → AMP consumes what is effectively 2 ~P bonds — the standard accounting is 1 ATP for charging); plus 1 GTP for EF-Tu/eEF1α (tRNA binding) and 1 GTP for EF-G/eEF2 (translocation). Total: 1 ATP + 2 GTP = 3 high-energy phosphate bonds.
  • The lac operon (catabolic) is induced by its substrate — default OFF. The trp operon (anabolic) is repressed by its product — default ON. This makes energetic sense because: Catabolic pathways break down nutrients; the enzymes are only needed when the substrate is available, so they should be OFF by default and turned ON when substrate appears. Anabolic pathways synthesize essential building blocks (like amino acids); the cell needs these continuously, so the pathway should be ON by default and turned OFF only when the product is already abundant in the environment — saving energy by not synthesizing what is readily available.

Keep learning

Ready to build on this? Continue to the next lesson.

Practice Biology 1

This lesson has no separate scored set. Practice draws from the subject’s question bank.

Study tools & related lessonsYou’ll learn to · Key vocabulary · Related

You’ll learn to

  • Describe transcription: identify the roles of promoter, RNA polymerase, and the template strand; summarize the stages of initiation, elongation, and termination
  • Compare transcription in prokaryotes and eukaryotes (location, RNA processing, polymerase types)
  • Explain eukaryotic RNA processing: 5′ cap, poly-A tail, intron/exon splicing, spliceosome function, and alternative splicing
  • Describe translation: decode the genetic code, identify start and stop codons, explain tRNA structure and aminoacyl-tRNA synthetase function
  • Diagram ribosomal structure and the A, P, and E sites; trace a polypeptide through initiation, elongation, and termination
  • Explain degeneracy (redundancy) of the genetic code and the wobble hypothesis
  • Outline gene regulation: compare prokaryotic operons (lac, trp) with eukaryotic mechanisms (chromatin, histone modification, DNA methylation, enhancers, silencers, transcription factors, combinatorial control)

Key vocabulary

Transcription
Synthesis of RNA from a DNA template by RNA polymerase
Promoter
DNA sequence where RNA polymerase binds to initiate transcription
Template strand
DNA strand read by RNA polymerase (3′ → 5′); complementary to the RNA transcript
Coding strand
Non-template DNA strand; has the same sequence as RNA (with T → U)
RNA polymerase
Enzyme that synthesizes RNA; does not require a primer
Sigma factor
Prokaryotic protein that directs RNA polymerase to specific promoters
Transcription bubble
Region of unwound DNA (~10–17 bp) where RNA synthesis occurs
Rho-independent termination
GC-rich hairpin + poly-U stretch causes transcript release
5′ cap
7-methylguanosine added to the 5′ end of eukaryotic mRNA; protects from degradation and facilitates translation
Poly-A tail
50–250 adenines added to eukaryotic mRNA 3′ end post-transcriptionally
Intron
Noncoding sequence removed from pre-mRNA during splicing
Exon
Coding sequence retained in mature mRNA
Spliceosome
Large RNA–protein complex (snRNPs) that catalyzes intron removal
snRNP
Small nuclear ribonucleoprotein; core component of the spliceosome (U1, U2, U4, U5, U6)
Alternative splicing
Production of multiple mRNA isoforms from a single gene by different exon combinations
Translation
Synthesis of a polypeptide chain using mRNA as a template
Codon
Three-nucleotide sequence in mRNA specifying one amino acid (or stop signal)
Start codon
AUG — codes for methionine; signals translation initiation
Stop (nonsense) codons
UAA, UAG, UGA — signal termination; no corresponding tRNA
Degeneracy (redundancy)
Most amino acids are specified by more than one codon
Wobble hypothesis
Relaxed base pairing at the third codon position allows one tRNA to recognize multiple codons
tRNA (transfer RNA)
Adaptor molecule carrying an amino acid; anticodon pairs with mRNA codon
Anticodon
Three-nucleotide tRNA sequence complementary to an mRNA codon
Aminoacyl-tRNA synthetase
Enzyme that attaches the correct amino acid to the correct tRNA
Ribosomal subunits (prokaryotic)
50S (large) + 30S (small) = 70S
Ribosomal subunits (eukaryotic)
60S (large) + 40S (small) = 80S
A site (aminoacyl)
Binds incoming aminoacyl-tRNA
P site (peptidyl)
Holds tRNA with the growing polypeptide chain
E site (exit)
Holds deacylated tRNA before it leaves the ribosome
Peptidyl transferase
rRNA ribozyme activity forming peptide bonds in the large subunit
Shine–Dalgarno sequence
Prokaryotic ribosome binding site upstream of AUG
Kozak sequence
Eukaryotic consensus sequence (GCCRCCAUGG) for translation initiation
Polysome (polyribosome)
Multiple ribosomes simultaneously translating one mRNA
Operon
Cluster of prokaryotic genes under control of a single promoter; transcribed as one mRNA
Operator
DNA sequence where a repressor binds to block transcription
Inducer
Molecule (e.g., allolactose) that inactivates a repressor → turns ON transcription
Corepressor
Molecule (e.g., tryptophan) that activates a repressor → turns OFF transcription
Enhancer
Distant DNA sequence that binds activator proteins and stimulates transcription
Silencer
DNA sequence that binds repressor proteins and reduces transcription
Transcription factor
Protein that binds DNA and regulates transcription (general or specific)
Combinatorial control
Regulation of gene expression by combinations of multiple transcription factors
Histone acetylation
Addition of acetyl groups to histone tails → opens chromatin → activates transcription
DNA methylation
Methylation of cytosine in CpG dinucleotides; generally represses transcription
Epigenetics
Heritable changes in gene expression without changes in DNA sequence
Euchromatin
Loosely packed, transcriptionally active chromatin
Heterochromatin
Tightly packed, transcriptionally silent chromatin

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.