Biology 1 · Genetics & Inheritance
Gene Expression: From DNA to Protein
On this page 4 sections
The college version
Core Explanation
Gene expression is the process by which information encoded in a gene — a specific sequence of DNA — is used to direct the synthesis of a functional gene product, most commonly a protein. In all organisms, gene expression proceeds through two major stages: Transcription Synthesis of RNA from a DNA template by RNA polymerase (DNA → RNA) and Translation Synthesis of a polypeptide chain using mRNA as a template (RNA → protein). This flow of genetic information is the central dogma of molecular biology, first articulated by Francis Crick in 1958.
DNA ──transcription──▶ RNA ──translation──▶ ProteinExceptions exist — some viruses use reverse transcription (RNA → DNA), and many RNA molecules (tRNA, rRNA, snRNA, microRNA) are functional without being translated — but the central dogma captures the fundamental information pathway for all cellular life.
Part I: Transcription — DNA-Directed RNA Synthesis
Transcription is the synthesis of an RNA molecule from a DNA template. The enzyme RNA polymerase Enzyme that synthesizes RNA; does not require a primer catalyzes this reaction. Only one of the two DNA strands — the Template strand DNA strand read by RNA polymerase (3′ → 5′); complementary to the RNA transcript (also called the antisense or noncoding strand) — serves as the template for RNA synthesis.
Key Players
| Component | Role |
|---|---|
| RNA polymerase | Enzyme that synthesizes RNA in the 5′ → 3′ direction by adding ribonucleoside triphosphates (NTPs) complementary to the DNA template. Does NOT require a primer. |
| Template strand | The DNA strand that RNA polymerase reads in the 3′ → 5′ direction. The synthesized RNA is complementary to this strand. |
| Coding strand | The non-template DNA strand. It has the same sequence as the RNA transcript (with T replaced by U). NOT used by RNA polymerase. |
| Promoter | A specific DNA sequence upstream of the gene where RNA polymerase binds to initiate transcription. Defines the transcription start site and which DNA strand will be the template. |
| Terminator | A DNA sequence that signals RNA polymerase to stop transcription and release the transcript. |
Critical distinction: The template strand is read 3′ → 5′ so the growing RNA chain is synthesized 5′ → 3′. The RNA transcript is complementary and antiparallel to the template strand. It is identical in sequence to the Coding strand Non-template DNA strand; has the same sequence as RNA (with T → U) (with U replacing T):
Coding strand: 5′ — ATGCCGTTAGAC — 3′
Template strand: 3′ — TACGGCAATCTG — 5′
RNA transcript: 5′ — AUGCCGUUAGAC — 3′Stages of Transcription
1. Initiation
In prokaryotes, RNA polymerase holoenzyme (core enzyme + Sigma factor Prokaryotic protein that directs RNA polymerase to specific promoters) recognizes and binds to the Promoter DNA sequence where RNA polymerase binds to initiate transcription. The promoter contains conserved sequences at approximately −10 (TATAAT, the Pribnow box) and −35 (TTGACA) relative to the transcription start site (+1). The sigma factor confers promoter specificity; once transcription begins, sigma dissociates.
In eukaryotes, initiation is far more complex. RNA polymerase II (which transcribes protein-coding genes) requires a suite of general transcription factors (TFIID, TFIIB, TFIIF, TFIIE, TFIIH, etc.) to assemble at the promoter. TFIID contains the TATA-binding protein (TBP) that recognizes the TATA box (consensus TATAAAA, located about −25 to −30). This assembly forms the preinitiation complex. TFIIH has helicase activity that unwinds DNA and kinase activity that phosphorylates RNA polymerase II, triggering promoter escape.
Key point: RNA polymerase unwinds approximately 10–17 base pairs of DNA, forming a Transcription bubble Region of unwound DNA (~10–17 bp) where RNA synthesis occurs. The first few NTPs are linked together without primer requirement (unlike DNA polymerase).
2. Elongation
RNA polymerase moves along the template strand 3′ → 5′, adding NTPs one at a time to the 3′ end of the growing RNA chain. Incoming NTPs pair with complementary bases on the template (A–U, G–C, T–A, C–G). A phosphodiester bond forms between the 3′-OH of the growing chain and the 5′-phosphate of the incoming NTP, releasing pyrophosphate (PPi).
The transcription bubble moves with the polymerase. DNA unwinds ahead and rewinds behind the bubble. The RNA transcript peels away from the template, and only about 8–9 nucleotides remain base-paired with DNA at any moment.
In eukaryotes, RNA polymerase II begins transcribing beyond the eventual 3′ end of the mRNA. Additional processing signals are encountered downstream.
3. Termination
Prokaryotic termination occurs by two mechanisms:
- Rho-independent (intrinsic) termination: The RNA transcript forms a GC-rich hairpin loop followed by a string of U residues. The hairpin causes RNA polymerase to pause, and the weak A–U base pairing (only 2 hydrogen bonds between A and U in the RNA–DNA hybrid, compared to 3 for G–C) allows the transcript to dissociate.
- Rho-dependent termination: The rho (ρ) protein, an ATP-dependent helicase, binds to a rut (rho utilization) site on the nascent RNA, translocates along the transcript, and unwinds the RNA–DNA hybrid when it catches up to the paused polymerase.
Eukaryotic termination: For protein-coding genes transcribed by RNA polymerase II, termination is coupled to RNA processing. After RNA polymerase transcribes past the polyadenylation signal (AAUAAA in the RNA), the transcript is cleaved ~10–35 nucleotides downstream. RNA polymerase II continues transcribing for hundreds of nucleotides before eventually dissociating; the trailing RNA is degraded. The cleaved pre-mRNA then undergoes processing.
Part II: Eukaryotic RNA Processing
In prokaryotes, transcription and translation are coupled: ribosomes begin translating mRNA while it is still being transcribed. In eukaryotes, the primary transcript (pre-mRNA) must undergo extensive processing in the nucleus before it can be exported to the cytoplasm for translation.
The Three Major Processing Events
1. 5′ Capping
Shortly after transcription begins (when the RNA is ~20–30 nucleotides long), a 7-methylguanosine cap is added to the 5′ end of the pre-mRNA via a 5′–5′ triphosphate linkage. The capping enzyme adds this modified guanine nucleotide in an unusual orientation.
Functions of the 5′ cap:
- Protects the mRNA from degradation by 5′ exonucleases
- Facilitates transport from the nucleus to the cytoplasm
- Recognized by the translation initiation complex (cap-binding proteins) for ribosome recruitment
2. 3′ Polyadenylation (Poly-A Tail)
After cleavage at the polyadenylation signal, the enzyme poly-A polymerase adds approximately 50–250 adenine nucleotides (the Poly-A tail 50–250 adenines added to eukaryotic mRNA 3′ end post-transcriptionally) to the new 3′ end. This addition does NOT use a DNA template — it is a post-transcriptional modification.
Functions of the poly-A tail:
- Protects the mRNA from 3′ exonuclease degradation
- Facilitates nuclear export
- Enhances translation initiation (poly-A binding proteins interact with cap-binding proteins, forming a closed-loop structure)
- The tail shortens over time in the cytoplasm; when it becomes too short, the mRNA is degraded — this is a mechanism for controlling mRNA lifespan
3. Splicing: Intron Removal
Most eukaryotic genes are interrupted by introns (intervening sequences) that do not code for protein. The coding sequences — exons — are expressed. Introns are removed and exons are joined together in a process called splicing, which occurs in the nucleus.
The Spliceosome Large RNA–protein complex (snRNPs) that catalyzes intron removal:
The spliceosome is a large RNA–protein complex that catalyzes splicing. It is composed of five small nuclear ribonucleoproteins (snRNPs, pronounced "snurps"): U1, U2, U4, U5, and U6. Each snRNP contains a small nuclear RNA (snRNA) and associated proteins. The snRNAs recognize the intron boundaries through complementary base pairing.
Splicing mechanism:
- U1 snRNP binds the 5′ splice site (GU dinucleotide at the intron's 5′ end).
- U2 snRNP binds the branch point adenine within the intron.
- U4/U6 and U5 snRNPs join, forming the active spliceosome.
- The 5′ end of the intron is cleaved and covalently linked to the branch point adenine, forming a lariat structure.
- The 3′ splice site (AG dinucleotide) is cleaved, and the two exons are ligated together.
- The intron lariat is released and degraded.
The GU–AG rule: nearly all introns begin with GU and end with AG. These conserved sequences are essential for splice site recognition.
What makes splicing remarkable: The spliceosome is a ribozyme — the catalytic activity comes from the snRNAs, not the proteins. This is a key piece of evidence for the RNA world hypothesis.
Alternative Splicing
Alternative splicing Production of multiple mRNA isoforms from a single gene by different exon combinations allows a single gene to produce multiple different mRNA transcripts (and thus multiple protein isoforms) by joining different combinations of exons. This is a primary explanation for why the ~20,000 protein-coding genes in the human genome can generate an estimated 100,000+ distinct proteins.
Types of alternative splicing include:
- Exon skipping (cassette exon): an exon is either included or excluded.
- Alternative 5′ or 3′ splice sites: the splice site boundary shifts, altering exon length.
- Intron retention: an intron is retained in the mature mRNA (common in plants, rarer in animals).
- Mutually exclusive exons: one of two exons is included, but not both.
Alternative splicing is regulated by RNA-binding proteins (splicing factors) that enhance (splicing enhancers) or repress (splicing silencers) splice site recognition. This regulation can be tissue-specific, developmental-stage-specific, or responsive to cellular signals.
Example: The Dscam gene in Drosophila has 38,016 possible alternative splice isoforms — more than the total number of genes in the fly genome. In humans, the troponin T gene produces different isoforms in cardiac and skeletal muscle through alternative splicing.
Pre-mRNA Processing Summary
Pre-mRNA: 5' ──[exon 1]──[==intron 1==]──[exon 2]──[==intron 2==]──[exon 3]── 3'
Processing: Add 5' cap → Splice out introns → Cleave and add poly-A tail
Mature mRNA: 5'-cap ──[exon 1]──[exon 2]──[exon 3]──AAAAAAAAAAA-3'Part III: Translation — RNA-Directed Protein Synthesis
Translation is the synthesis of a polypeptide chain using the information encoded in mRNA. It occurs on ribosomes in the cytoplasm (or on the rough ER). The language of nucleic acids (nucleotides) is translated into the language of proteins (amino acids).
The Genetic Code
The genetic code is the set of rules by which nucleotide triplets (codons) specify amino acids. Each codon consists of three consecutive nucleotides in the mRNA (read 5′ → 3′).
Fundamental properties of the genetic code:
| Property | Description |
|---|---|
| Triplet code | Three nucleotides = one codon = one amino acid (or stop signal) |
| Non-overlapping | Nucleotides are read in consecutive, non-overlapping groups of three |
| Commaless | No punctuation between codons; the reading frame is maintained continuously |
| Degenerate (redundant) | Most amino acids are specified by more than one codon (see below) |
| Unambiguous | Each codon specifies only ONE amino acid (no codon codes for two different amino acids) |
| Nearly universal | Shared by almost all organisms (minor exceptions in mitochondria and some protists) |
| Start codon | AUG — codes for methionine; also serves as the initiation signal |
| Stop codons | UAA, UAG, UGA — do not code for any amino acid; signal termination |
Of the 64 possible codons: 61 code for amino acids, and 3 are stop codons.
The Genetic Code Table
Second base
U C A G
┌─────────┬─────────┬─────────┬─────────┐
U │ UUU Phe │ UCU Ser │ UAU Tyr │ UGU Cys │ U
│ UUC Phe │ UCC Ser │ UAC Tyr │ UGC Cys │ C
│ UUA Leu │ UCA Ser │ UAA STOP│ UGA STOP│ A
│ UUG Leu │ UCG Ser │ UAG STOP│ UGG Trp │ G
├─────────┼─────────┼─────────┼─────────┤
C │ CUU Leu │ CCU Pro │ CAU His │ CGU Arg │ U
F │ CUC Leu │ CCC Pro │ CAC His │ CGC Arg │ C
i │ CUA Leu │ CCA Pro │ CAA Gln │ CGA Arg │ A
r │ CUG Leu │ CCG Pro │ CAG Gln │ CGG Arg │ G
s ├─────────┼─────────┼─────────┼─────────┤
t A │ AUU Ile │ ACU Thr │ AAU Asn │ AGU Ser │ U
│ AUC Ile │ ACC Thr │ AAC Asn │ AGC Ser │ C
b │ AUA Ile │ ACA Thr │ AAA Lys │ AGA Arg │ A
a │ AUG Met │ ACG Thr │ AAG Lys │ AGG Arg │ G
s ├─────────┼─────────┼─────────┼─────────┤
e G │ GUU Val │ GCU Ala │ GAU Asp │ GGU Gly │ U
│ GUC Val │ GCC Ala │ GAC Asp │ GGC Gly │ C
│ GUA Val │ GCA Ala │ GAA Glu │ GGA Gly │ A
│ GUG Val │ GCG Ala │ GAG Glu │ GGG Gly │ G
└─────────┴─────────┴─────────┴─────────┘Degeneracy and the Wobble Hypothesis
The code is degenerate: 18 of the 20 amino acids are encoded by more than one codon. Methionine and tryptophan are the only exceptions, each having a single codon (AUG and UGG, respectively).
Codons for the same amino acid tend to differ only in the third base position — often called the "wobble" position. The Wobble hypothesis Relaxed base pairing at the third codon position allows one tRNA to recognize multiple codons (proposed by Francis Crick in 1966) explains why fewer tRNAs are needed than there are codons: the base pairing between the third position of the codon (mRNA) and the first position of the anticodon (tRNA) is less stringent — nonstandard base pairing (e.g., G–U, inosine–U/C/A) is permitted.
This means a single tRNA can recognize multiple codons for the same amino acid. For example, there are six codons for leucine but far fewer leucine-specific tRNAs, because wobble pairing allows one tRNA anticodon to pair with multiple leucine codons.
Why degeneracy matters: It buffers against the effects of point mutations. A change in the third codon position often produces a silent mutation — the same amino acid is still incorporated.
Transfer RNA (tRNA): The Adaptor Molecule
tRNA is the molecular adaptor that translates the nucleic acid language of mRNA into the amino acid language of proteins. Each tRNA molecule has two critical functional regions:
- Anticodon: A three-nucleotide sequence that base-pairs (antiparallel) with a complementary mRNA codon.
- Acceptor stem (3′ end): The sequence CCA at the 3′ terminus, where the corresponding amino acid is covalently attached.
tRNA folds into a characteristic cloverleaf secondary structure (in 2D) and an L-shaped tertiary structure (in 3D). This precise shape is essential for fitting into the ribosome's A and P sites.
Aminoacyl-tRNA Synthetase
Each of the 20 amino acids has its own aminoacyl-tRNA synthetase — an enzyme that catalyzes the attachment of the correct amino acid to the correct tRNA in a two-step reaction:
Amino acid + ATP → Aminoacyl-AMP + PPi
Aminoacyl-AMP + tRNA → Aminoacyl-tRNA + AMPThe charged tRNA is now called an aminoacyl-tRNA. The high-energy bond between the amino acid and the tRNA provides the energy for peptide bond formation later.
Proofreading: Aminoacyl-tRNA synthetases have editing sites that hydrolyze incorrectly attached amino acids. This double-check mechanism is essential — an error rate of even 1 in 10,000 at this step would introduce errors into nearly every protein. The overall accuracy of aminoacyl-tRNA synthetases exceeds 99.99%.
The Ribosome: The Protein Synthesis Factory
Ribosomes are large ribonucleoprotein complexes composed of two subunits.
| Prokaryotic Ribosome | Eukaryotic Ribosome | |
|---|---|---|
| Overall | 70S | 80S |
| Large subunit | 50S (23S rRNA + 5S rRNA + ~31 proteins) | 60S (28S rRNA + 5.8S rRNA + 5S rRNA + ~49 proteins) |
| Small subunit | 30S (16S rRNA + ~21 proteins) | 40S (18S rRNA + ~33 proteins) |
Three tRNA binding sites are formed at the interface of the two subunits:
| Site | Full Name | Function |
|---|---|---|
| A site | Aminoacyl site | Accepts the incoming aminoacyl-tRNA (except initiator tRNA) |
| P site | Peptidyl site | Holds the tRNA carrying the growing polypeptide chain |
| E site | Exit site | Holds the deacylated tRNA (now without an amino acid) before it exits the ribosome |
Critical: The peptidyl transferase activity that forms peptide bonds is catalyzed by the 23S rRNA (prokaryotes) or 28S rRNA (eukaryotes) in the large subunit — not by any ribosomal protein. The ribosome is a ribozyme. This discovery (Steitz, Yonath, Ramakrishnan — Nobel Prize in Chemistry, 2009) provided powerful evidence for the RNA world hypothesis.
Stages of Translation
1. Initiation
Prokaryotic initiation:
- The 30S small subunit binds to the Shine–Dalgarno sequence (AGGAGG) in the mRNA, located ~5–10 nucleotides upstream of the start codon (AUG). This sequence is complementary to the 3′ end of 16S rRNA, positioning the start codon in the P site.
- The initiator tRNA (carrying N-formylmethionine, fMet) binds to the AUG start codon in the P site with the help of initiation factors (IF1, IF2, IF3).
- The 50S large subunit joins, forming the 70S initiation complex. Initiation factors are released, and GTP is hydrolyzed.
Eukaryotic initiation (simplified):
- The 40S small subunit, bound to initiation factors and the initiator tRNA (carrying methionine, not formylated), scans the mRNA from the 5′ cap until it encounters an AUG in a favorable sequence context (Kozak sequence: GCCRCCAUGG, where R = purine).
- Upon AUG recognition, the 60S large subunit joins, forming the 80S initiation complex.
- Note: eukaryotic initiation is FAR more complex — over a dozen initiation factors are involved (eIFs).
Key difference: Prokaryotes use a Shine–Dalgarno sequence for ribosome binding; eukaryotes use cap-dependent scanning.
2. Elongation
Elongation is a cyclic three-step process (repeated for each amino acid added):
Step 1 — Codon recognition (tRNA binding): An aminoacyl-tRNA with an anticodon complementary to the codon in the A site enters the A site. This process requires elongation factor EF-Tu (prokaryotes) or eEF1α (eukaryotes) and GTP hydrolysis for proofreading.
Step 2 — Peptide bond formation (transpeptidation): The peptidyl transferase center (the rRNA of the large subunit) catalyzes the formation of a peptide bond between the amino group of the incoming amino acid (in the A site) and the carboxyl end of the growing chain (attached to the tRNA in the P site). The polypeptide chain is transferred from the tRNA in the P site to the tRNA in the A site. No additional energy input is needed — the energy comes from the high-energy aminoacyl-tRNA bond.
Step 3 — Translocation: The ribosome moves (translocates) one codon toward the 3′ end of the mRNA. This requires elongation factor EF-G (prokaryotes) or eEF2 (eukaryotes) and GTP hydrolysis. After translocation:
- The deacylated tRNA (now uncharged) moves from the P site to the E site, then exits.
- The peptidyl-tRNA (with the growing chain) moves from the A site to the P site.
- The A site is now empty and aligned with the next codon, ready for the next aminoacyl-tRNA.
Energy cost per amino acid added: 2 GTP (one for tRNA binding, one for translocation) + 1 ATP equivalent (from the aminoacyl-tRNA charging reaction) = 3 high-energy phosphate bonds per peptide bond.
3. Termination
Termination occurs when a stop codon (UAA, UAG, or UGA) enters the A site:
- No tRNA recognizes stop codons. Instead, a protein release factor (RF1 or RF2 in prokaryotes; eRF1 in eukaryotes) binds the stop codon in the A site. Release factors have a shape that mimics tRNA.
- The release factor triggers the peptidyl transferase center to hydrolyze the bond linking the polypeptide chain to the tRNA in the P site (adding a water molecule instead of an amino acid).
- The polypeptide chain is released from the ribosome.
- The ribosomal subunits, mRNA, tRNA, and release factors dissociate (aided by GTP hydrolysis and ribosome recycling factors).
Polysomes (polyribosomes): Multiple ribosomes can simultaneously translate a single mRNA molecule, each at a different stage of translation. A typical eukaryotic mRNA might have 5–10 ribosomes translating it at once. This allows a single mRNA to produce many protein copies before it is degraded.
Translation Summary Table
| Stage | Key Events |
|---|---|
| Initiation | Small subunit binds mRNA → initiator tRNA pairs with AUG at P site → large subunit joins |
| Elongation | Aminoacyl-tRNA enters A site → peptide bond forms (P → A) → translocation (A → P, P → E) → repeat |
| Termination | Stop codon enters A site → release factor binds → polypeptide released → ribosome dissociates |
Part IV: Gene Regulation — An Overview
Not all genes are expressed at all times. Cells regulate which genes are transcribed and translated, controlling both the timing and the amount of gene expression. This regulation is essential for development, differentiation, homeostasis, and response to environmental signals.
Prokaryotic Gene Regulation: Operons
Prokaryotes often organize functionally related genes into operons — clusters of genes under the control of a single promoter that are transcribed together as a single polycistronic mRNA.
The lac Operon: Inducible (Catabolic) System
The lac operon in E. coli contains three structural genes (lacZ, lacY, lacA) encoding enzymes for lactose metabolism: β-galactosidase (lacZ), lactose permease (lacY), and thiogalactoside transacetylase (lacA). These genes are transcribed as a single mRNA.
Regulatory components:
- Promoter (P): Binding site for RNA polymerase.
- Operator (O): DNA sequence between the promoter and structural genes; the binding site for the lac repressor.
- lacI gene: Located upstream of the operon; constitutively expressed. It encodes the lac repressor protein.
- CAP binding site: Upstream of the promoter; binding site for catabolite activator protein (CAP).
Regulation logic:
| Condition | Repressor | CAP | Transcription |
|---|---|---|---|
| No lactose, glucose present | Bound to operator | Inactive (no cAMP) | OFF |
| Lactose, glucose present | Inactive (bound to allolactose) | Inactive (no cAMP) | Low (basal) |
| Lactose, no glucose | Inactive (bound to allolactose) | Active (cAMP bound, CAP binds) | HIGH |
| No lactose, no glucose | Bound to operator | Active (CAP bound) | OFF (repressor blocks) |
The dual control ensures efficiency: The operon is only fully ON when lactose is available AND glucose is absent. Glucose is the preferred energy source — if glucose is present, it is energetically wasteful to produce lactose-metabolizing enzymes.
How allolactose deactivates the repressor: A small amount of lactose is converted to allolactose (the inducer) by existing β-galactosidase. Allolactose binds to the lac repressor, causing a conformational change that prevents the repressor from binding the operator. This is an inducible system — the substrate (lactose) induces enzyme synthesis.
The trp Operon: Repressible (Anabolic) System
The trp operon encodes five enzymes for tryptophan biosynthesis.
Regulatory logic: Unlike the lac operon (which is induced by its substrate), the trp operon is repressed by its product. When tryptophan is abundant, the cell does not need to synthesize more.
Mechanism:
- The trp repressor is produced in an inactive form by a separate regulatory gene (trpR).
- When tryptophan levels are high, tryptophan acts as a corepressor — it binds to the trp repressor, activating it.
- The active repressor–tryptophan complex binds to the operator, blocking transcription.
| Condition | Repressor | Transcription |
|---|---|---|
| Low tryptophan | Inactive (no corepressor bound) | ON |
| High tryptophan | Active (tryptophan corepressor bound) | OFF |
Attenuation: The trp operon also uses a second regulatory mechanism called attenuation, which couples transcription and translation. When tryptophan is abundant, ribosomes translate the leader peptide quickly, causing the formation of a transcription terminator hairpin in the mRNA — premature termination. Attenuation allows fine-tuning beyond simple on/off control.
Summary: Inducible vs Repressible Operons
| Feature | lac Operon | trp Operon |
|---|---|---|
| Pathway type | Catabolic (breakdown) | Anabolic (synthesis) |
| Regulation type | Inducible | Repressible |
| Default state | OFF | ON |
| Effector molecule | Allolactose (inducer) | Tryptophan (corepressor) |
| Effector effect | Inactivates repressor → ON | Activates repressor → OFF |
| Responds to | Presence of substrate | Abundance of product |
Exam tip: Inducible systems (lac, ara) control catabolic pathways — turn ON when substrate is present. Repressible systems (trp, his) control anabolic pathways — turn OFF when product is abundant. This makes energetic sense.
Eukaryotic Gene Regulation
Eukaryotic gene regulation is vastly more complex than prokaryotic regulation. Every step in the gene expression pathway can be controlled — from chromatin accessibility through transcription, RNA processing, mRNA export, translation, and protein degradation. Below is an overview of the major mechanisms, focused primarily on transcriptional regulation.
1. Chromatin Structure and Accessibility
In eukaryotes, DNA is packaged with histone proteins into chromatin. The default state is "packed up" and inaccessible. Gene activation requires chromatin remodeling to expose promoter regions.
Euchromatin: Loosely packed, transcriptionally active. \ Heterochromatin: Tightly packed, transcriptionally silent. (Includes constitutive heterochromatin — always silent, such as centromeres and telomeres — and facultative heterochromatin, which can be reactivated depending on the cell type or developmental stage.)
Chromatin remodeling complexes (e.g., SWI/SNF) use ATP hydrolysis to slide or eject nucleosomes, exposing DNA to transcription factors and RNA polymerase.
2. Histone Modification
The N-terminal tails of histone proteins protrude from the nucleosome core and are subject to various covalent modifications that affect chromatin structure and gene expression:
| Modification | Effect on Transcription | Key Enzymes |
|---|---|---|
| Histone acetylation (adding acetyl groups to lysine) | Activates transcription — neutralizes positive charge on lysine, weakening histone–DNA interaction (opens chromatin) | Histone acetyltransferases (HATs); removed by histone deacetylases (HDACs) |
| Histone methylation (adding methyl groups to lysine or arginine) | Context-dependent: can activate OR repress depending on which residue is methylated | Histone methyltransferases (HMTs); removed by histone demethylases |
| Histone phosphorylation | Often associated with chromatin condensation during mitosis and transcriptional activation at specific genes | Kinases; removed by phosphatases |
The histone code hypothesis: Combinations of histone modifications create a "code" that is read by other proteins to determine the transcriptional state of a gene. For example, H3K4me3 (trimethylation of lysine 4 on histone H3) is strongly associated with active promoters, while H3K27me3 is associated with gene silencing.
3. DNA Methylation
DNA methylation — the addition of a methyl group to the 5′ position of cytosine, typically in CpG dinucleotides — is generally associated with transcriptional repression:
- Methylated CpG sequences in promoter regions recruit methyl-CpG-binding proteins, which in turn recruit histone deacetylases (HDACs) and other repressive chromatin modifiers.
- CpG islands (GC-rich regions near promoters) are usually unmethylated in active genes and methylated in silenced genes.
DNA methylation patterns are heritable through cell division (maintained by DNA methyltransferases like DNMT1), providing a mechanism for epigenetic inheritance — stable changes in gene expression that do NOT involve changes in the DNA sequence itself.
Examples: X-chromosome inactivation in female mammals and genomic imprinting both involve DNA methylation. Aberrant DNA methylation is also a hallmark of many cancers — tumor suppressor genes are often hypermethylated (silenced), while oncogenes may be hypomethylated (activated).
4. Enhancers and Silencers
Enhancers are DNA sequences (typically 50–1,500 bp) that can be located thousands of base pairs away from the promoter — upstream, downstream, within introns, or even on a different chromosome (in trans). They bind activator proteins (specific transcription factors) and increase transcription of associated genes.
Silencers are analogous DNA sequences that bind repressor proteins and decrease transcription.
How do enhancers work over long distances? DNA looping brings enhancer-bound activator proteins into physical contact with the promoter-bound transcription machinery. The intervening DNA loops out, and mediator complexes and cohesin help bridge and stabilize the interaction.
Key properties of enhancers:
- They work in an orientation-independent manner (can be inverted and still function).
- They can act at great distances (some enhancers are >1 Mb from their target gene).
- They are cell-type-specific — a given enhancer is active only in cell types where the appropriate activator proteins are present.
- Insulators (boundary elements) block enhancer–promoter interactions, defining regulatory domains and preventing inappropriate gene activation.
5. Transcription Factors
Transcription factors are proteins that bind specific DNA sequences and regulate transcription.
General (basal) transcription factors: Required for transcription by RNA polymerase II at ALL promoters (e.g., TFIID, TFIIB, TFIIH). They are part of the basic transcriptional machinery — necessary but not sufficient for high-level, regulated expression.
Specific transcription factors (regulatory): Bind to enhancers, silencers, or promoter-proximal elements and regulate transcription of SPECIFIC genes. They contain at least two functional domains:
- DNA-binding domain: Recognizes a specific DNA sequence (e.g., helix-turn-helix, zinc finger, leucine zipper, helix-loop-helix motifs).
- Activation domain (or repression domain): Interacts with other proteins (mediator complex, chromatin modifiers, general transcription factors) to stimulate or repress transcription.
6. Combinatorial Control
A single gene is typically regulated by many different transcription factors acting in combination. Conversely, a single transcription factor can regulate many different genes. This combinatorial control allows a limited number of transcription factors (humans have ~1,600) to regulate >20,000 genes in precise, tissue-specific, and signal-responsive patterns.
How combinatorial control works:
- Multiple activator proteins must bind simultaneously for a gene to be transcribed (AND logic).
- Different combinations of transcription factors activate different target genes in different cell types.
- The same transcription factor can activate one gene and repress another, depending on its binding partners at each locus.
Example: The human β-globin gene is regulated by at least five different transcription factors binding to its enhancer and promoter. The specific combination of factors present in erythroid precursor cells (but not in other cell types) drives high-level expression — even though some of those factors are present in other tissues, the full complement is only present in red blood cell precursors.
Prokaryotic vs Eukaryotic Gene Regulation at a Glance
| Feature | Prokaryotes | Eukaryotes |
|---|---|---|
| Organization | Operons (polycistronic mRNAs) | Individual genes (mostly monocistronic) |
| Chromatin | No (naked DNA, though nucleoid-associated proteins exist) | Yes — chromatin structure is a major regulatory layer |
| Primary control level | Transcription initiation (operators, repressors, activators) | Transcription initiation (chromatin, enhancers, TFs) but also RNA processing, export, mRNA stability, translation |
| Regulatory proteins | Repressors and activators (simple) | Hundreds of transcription factors + coactivators, corepressors, chromatin modifiers |
| DNA looping | CAP bends DNA; AraC loops DNA | Extensive long-range enhancer–promoter looping |
| Post-transcriptional | Minimal (mRNA is short-lived) | Extensive (alternative splicing, RNA editing, RNAi, mRNA localization, differential stability) |
Biological / Medical Relevance
- Antibiotics targeting transcription and translation: Rifampin binds bacterial RNA polymerase (treats tuberculosis). Tetracycline blocks tRNA binding to the A site; chloramphenicol inhibits peptidyl transferase; erythromycin blocks translocation. These selectively target prokaryotic ribosomes (70S), minimizing host toxicity. Streptomycin causes misreading of the genetic code.
- Splicing defects and disease: ~15% of human genetic diseases are caused by mutations that disrupt splicing. Spinal muscular atrophy (SMA) results from mutations affecting SMN2 splicing. The drug nusinersen (Spinraza) is an antisense oligonucleotide that corrects SMN2 splicing.
- Cancer and gene regulation: Aberrant DNA methylation (hypermethylation of tumor suppressors, hypomethylation of oncogenes), histone modification defects, and mutations in chromatin remodeling complexes (e.g., SWI/SNF) are hallmarks of cancer.
- Epigenetics and development: Genomic imprinting disorders (Prader-Willi, Angelman syndromes) result from defects in DNA methylation at imprinted loci. X-inactivation in female mammals is an epigenetic phenomenon governed by the lncRNA XIST.
- CRISPR-Cas9 gene editing: Engineered transcription factors (CRISPR activation, CRISPR interference) can be targeted to specific promoters or enhancers to activate or repress endogenous genes — directly leveraging the gene regulation principles described here.
- COVID-19 mRNA vaccines: The Pfizer-BioNTech and Moderna vaccines deliver synthetic mRNA with a 5′ cap and poly-A tail, optimized codons, and modified nucleosides — a direct clinical application of the gene expression machinery.
Common Misconceptions and Exam Traps
- Exam trap: The template strand is read 3′ → 5′, and RNA is synthesized 5′ → 3′. Students routinely get this backward. Think: polymerases ALWAYS synthesize 5′ → 3′.
- Exam trap: "The coding strand is the template." WRONG. The coding strand has the SAME sequence as RNA (T → U); the template strand is the one that RNA polymerase reads.
- Misconception: "RNA polymerase needs a primer." RNA polymerase does NOT require a primer, unlike DNA polymerase. This is a classic exam distinction.
- Exam trap: "Ribosomal proteins catalyze peptide bond formation." WRONG. The ribosome is a ribozyme — the 23S/28S rRNA catalyzes peptide bond formation. This is a Nobel Prize-level concept that exams love.
- Misconception: "The lac operon is always ON when lactose is present." False. With glucose present, cAMP is low, CAP is inactive, and transcription is only at a basal (low) level. The lac operon requires BOTH lactose present AND glucose absent for full activation.
- Exam trap: Confusing lac (inducible → substrate turns it ON) with trp (repressible → product turns it OFF). An easy mnemonic: inducible = catabolic (breakdown); repressible = anabolic (synthesis). Makes energetic sense.
- Misconception: "Degeneracy means the genetic code is ambiguous." WRONG. The code is degenerate (multiple codons per amino acid) but UNAMBIGUOUS (each codon codes for only ONE amino acid). These are distinct properties — exams test this distinction.
- Exam trap: "The stop codon is recognized by a special tRNA." WRONG. Stop codons are recognized by protein release factors (RF1/RF2/eRF1), NOT by tRNAs.
- Misconception: "Alternative splicing means the same protein is made in different ways." Wrong direction — alternative splicing means different mRNA transcripts (and thus DIFFERENT proteins) are made from the SAME gene.
- Misconception: "DNA methylation always silences genes." While generally repressive, DNA methylation in gene bodies (as opposed to promoters) can sometimes be associated with active transcription. The promoter context matters.

Eli explains
The same idea, in plain words
Explain it like I’m 10
Your DNA is a giant recipe book, but it stays locked in the nucleus (the "library" of the cell). When the cell needs to make a protein, it can't take the original recipe out — that's too risky. So it makes a copy of just the page it needs. That's transcription — copying a gene from DNA into messenger RNA. Before the copy leaves the nucleus, the cell edits it: it adds a protective cap to the front, a tail of extra letters to the back, and cuts out the "nonsense" parts (introns), gluing the useful parts (exons) together. The finished message heads out to the cytoplasm, where ribosomes — tiny machines made of RNA and protein — read the message three letters at a time. Each three-letter word (codon) means one amino acid. Special carriers called tRNA bring the right amino acid for each codon, and the ribosome links them into a chain. That chain folds up into a working protein. The cell also has ways to decide WHICH recipes to copy, and how many times — that's gene regulation. Prokaryotic cells use simple switches (operons); eukaryotic cells have a whole control room full of regulators, chemical tags on DNA and histones, and remote on/off switches called enhancers.
Key takeaways
- The central dogma: DNA → RNA → Protein. Transcription produces RNA; translation produces protein.
- RNA polymerase reads the template strand 3′ → 5′; RNA is synthesized 5′ → 3′. No primer required.
- Promoters define the transcription start site and DNA strand to be transcribed. Eukaryotes use TATA box + general transcription factors; prokaryotes use −10 and −35 consensus sequences.
- Eukaryotic pre-mRNA processing: 5′ cap (protection + ribosome recruitment), splicing (intron removal by spliceosome), poly-A tail (protection + export).
- The spliceosome is a ribozyme — catalytic activity comes from snRNA, not protein.
- Alternative splicing: one gene → multiple protein isoforms. Explains human proteome complexity.
- Genetic code: triplet, non-overlapping, degenerate, unambiguous, nearly universal.
- Start codon = AUG (Met); stop codons = UAA, UAG, UGA.
- Wobble: relaxed third-position base pairing allows fewer tRNAs to cover all codons.
- tRNA is the adaptor; aminoacyl-tRNA synthetase charges tRNA with correct amino acid (proofreads).
- Ribosome A site (incoming tRNA), P site (peptidyl-tRNA), E site (exit). Peptide bonds are catalyzed by rRNA (ribozyme).
- Translation: initiation (small subunit + initiator tRNA at AUG → large subunit joins), elongation (codon recognition → peptide bond → translocation × N), termination (release factor at stop codon).
- Energy cost: ~3 high-energy phosphate bonds per peptide bond formed.
- lac operon: inducible, catabolic. Default OFF; induced by allolactose (inactivates repressor). Also subject to positive control by CAP (glucose sensing).
- trp operon: repressible, anabolic. Default ON; repressed by tryptophan (corepressor activates repressor).
- Eukaryotic regulation is multilayered: chromatin remodeling, histone modifications (acetylation = ON), DNA methylation (generally OFF), enhancers/silencers, combinatorial TF control.
- Enhancers act at a distance, orientation-independent, and are cell-type-specific.
- ---
- Transcription: RNA polymerase binds promoter → reads template strand 3′ → 5′ → synthesizes RNA 5′ → 3′ (initiation → elongation → termination). No primer needed.
- Eukaryotic RNA processing: 5′ cap (protection) + intron splicing (spliceosome — a ribozyme) + poly-A tail (protection + export).
- Genetic code: Triplet, degenerate, unambiguous, nearly universal. AUG = Met (start); UAA/UAG/UGA = stop.
- tRNA: Anticodon pairs with codon; aminoacyl-tRNA synthetase charges tRNA (proofreads). Wobble at 3rd position.
- Ribosome: A site (incoming), P site (peptide), E site (exit). Peptidyl transferase = rRNA (ribozyme).
- Translation: Initiation (small subunit + initiator tRNA → large subunit) → elongation (3-step cycle: binding, peptide bond, translocation) → termination (release factor at stop codon).
- lac operon: Inducible (substrate → ON). trp operon: Repressible (product → OFF).
- Eukaryotic regulation: Chromatin (euchromatin/heterochromatin), histone acetylation (ON), DNA methylation (OFF), enhancers/silencers, combinatorial TF control.
- ---
- A gene has the following template strand sequence (partial): 3′ — TACGATC — 5′. What is the sequence of the RNA transcript produced from this region?
- Why does the lac operon NOT produce high levels of β-galactosidase when both glucose and lactose are present?
- A mutation changes a codon from GAA to GAG. Both code for glutamic acid. What type of mutation is this, and why is it possible?
- How many high-energy phosphate bonds are consumed for each amino acid added to a growing polypeptide chain during elongation (including the energy cost of tRNA charging)?
- Contrast the default state and regulatory logic of the lac operon with the trp operon. Why does it make energetic sense that catabolic operons are inducible and anabolic operons are repressible?
- ---
- Template strand (read 3′ → 5′): 3′ — TACGATC — 5′. RNA is synthesized 5′ → 3′, complementary to the template (A → U, T → A, C → G, G → C). The RNA transcript is: 5′ — AUGCUAG — 3′. (This is the same sequence as the coding strand, with T replaced by U.)
- The lac operon requires BOTH the absence of the repressor AND the presence of active CAP for high-level transcription. With glucose present, intracellular cAMP levels are low, so CAP remains inactive and cannot bind its site to recruit RNA polymerase. The lac repressor is inactivated by allolactose, so there is low (basal) transcription, but without CAP-mediated positive control, transcription is NOT fully activated. The cell conserves energy by preferentially using glucose.
- This is a silent mutation — a nucleotide change that does not alter the amino acid sequence of the protein. It is possible because of the degeneracy (redundancy) of the genetic code: both GAA and GAG specify glutamic acid. The change occurred at the third (wobble) codon position, where base-pairing is less stringent.
- Three high-energy phosphate bonds per amino acid: one from ATP during tRNA charging (aminoacyl-tRNA synthetase: amino acid + ATP → aminoacyl-AMP + PPi; this is equivalent to 2 high-energy bonds, but the conversion of ATP → AMP consumes what is effectively 2 ~P bonds — the standard accounting is 1 ATP for charging); plus 1 GTP for EF-Tu/eEF1α (tRNA binding) and 1 GTP for EF-G/eEF2 (translocation). Total: 1 ATP + 2 GTP = 3 high-energy phosphate bonds.
- The lac operon (catabolic) is induced by its substrate — default OFF. The trp operon (anabolic) is repressed by its product — default ON. This makes energetic sense because: Catabolic pathways break down nutrients; the enzymes are only needed when the substrate is available, so they should be OFF by default and turned ON when substrate appears. Anabolic pathways synthesize essential building blocks (like amino acids); the cell needs these continuously, so the pathway should be ON by default and turned OFF only when the product is already abundant in the environment — saving energy by not synthesizing what is readily available.
Study tools & related lessonsYou’ll learn to · Key vocabulary · Related
You’ll learn to
- Describe transcription: identify the roles of promoter, RNA polymerase, and the template strand; summarize the stages of initiation, elongation, and termination
- Compare transcription in prokaryotes and eukaryotes (location, RNA processing, polymerase types)
- Explain eukaryotic RNA processing: 5′ cap, poly-A tail, intron/exon splicing, spliceosome function, and alternative splicing
- Describe translation: decode the genetic code, identify start and stop codons, explain tRNA structure and aminoacyl-tRNA synthetase function
- Diagram ribosomal structure and the A, P, and E sites; trace a polypeptide through initiation, elongation, and termination
- Explain degeneracy (redundancy) of the genetic code and the wobble hypothesis
- Outline gene regulation: compare prokaryotic operons (lac, trp) with eukaryotic mechanisms (chromatin, histone modification, DNA methylation, enhancers, silencers, transcription factors, combinatorial control)
Key vocabulary
- Transcription
- Synthesis of RNA from a DNA template by RNA polymerase
- Promoter
- DNA sequence where RNA polymerase binds to initiate transcription
- Template strand
- DNA strand read by RNA polymerase (3′ → 5′); complementary to the RNA transcript
- Coding strand
- Non-template DNA strand; has the same sequence as RNA (with T → U)
- RNA polymerase
- Enzyme that synthesizes RNA; does not require a primer
- Sigma factor
- Prokaryotic protein that directs RNA polymerase to specific promoters
- Transcription bubble
- Region of unwound DNA (~10–17 bp) where RNA synthesis occurs
- Rho-independent termination
- GC-rich hairpin + poly-U stretch causes transcript release
- 5′ cap
- 7-methylguanosine added to the 5′ end of eukaryotic mRNA; protects from degradation and facilitates translation
- Poly-A tail
- 50–250 adenines added to eukaryotic mRNA 3′ end post-transcriptionally
- Intron
- Noncoding sequence removed from pre-mRNA during splicing
- Exon
- Coding sequence retained in mature mRNA
- Spliceosome
- Large RNA–protein complex (snRNPs) that catalyzes intron removal
- snRNP
- Small nuclear ribonucleoprotein; core component of the spliceosome (U1, U2, U4, U5, U6)
- Alternative splicing
- Production of multiple mRNA isoforms from a single gene by different exon combinations
- Translation
- Synthesis of a polypeptide chain using mRNA as a template
- Codon
- Three-nucleotide sequence in mRNA specifying one amino acid (or stop signal)
- Start codon
- AUG — codes for methionine; signals translation initiation
- Stop (nonsense) codons
- UAA, UAG, UGA — signal termination; no corresponding tRNA
- Degeneracy (redundancy)
- Most amino acids are specified by more than one codon
- Wobble hypothesis
- Relaxed base pairing at the third codon position allows one tRNA to recognize multiple codons
- tRNA (transfer RNA)
- Adaptor molecule carrying an amino acid; anticodon pairs with mRNA codon
- Anticodon
- Three-nucleotide tRNA sequence complementary to an mRNA codon
- Aminoacyl-tRNA synthetase
- Enzyme that attaches the correct amino acid to the correct tRNA
- Ribosomal subunits (prokaryotic)
- 50S (large) + 30S (small) = 70S
- Ribosomal subunits (eukaryotic)
- 60S (large) + 40S (small) = 80S
- A site (aminoacyl)
- Binds incoming aminoacyl-tRNA
- P site (peptidyl)
- Holds tRNA with the growing polypeptide chain
- E site (exit)
- Holds deacylated tRNA before it leaves the ribosome
- Peptidyl transferase
- rRNA ribozyme activity forming peptide bonds in the large subunit
- Shine–Dalgarno sequence
- Prokaryotic ribosome binding site upstream of AUG
- Kozak sequence
- Eukaryotic consensus sequence (GCCRCCAUGG) for translation initiation
- Polysome (polyribosome)
- Multiple ribosomes simultaneously translating one mRNA
- Operon
- Cluster of prokaryotic genes under control of a single promoter; transcribed as one mRNA
- Operator
- DNA sequence where a repressor binds to block transcription
- Inducer
- Molecule (e.g., allolactose) that inactivates a repressor → turns ON transcription
- Corepressor
- Molecule (e.g., tryptophan) that activates a repressor → turns OFF transcription
- Enhancer
- Distant DNA sequence that binds activator proteins and stimulates transcription
- Silencer
- DNA sequence that binds repressor proteins and reduces transcription
- Transcription factor
- Protein that binds DNA and regulates transcription (general or specific)
- Combinatorial control
- Regulation of gene expression by combinations of multiple transcription factors
- Histone acetylation
- Addition of acetyl groups to histone tails → opens chromatin → activates transcription
- DNA methylation
- Methylation of cytosine in CpG dinucleotides; generally represses transcription
- Epigenetics
- Heritable changes in gene expression without changes in DNA sequence
- Euchromatin
- Loosely packed, transcriptionally active chromatin
- Heterochromatin
- Tightly packed, transcriptionally silent chromatin
Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.
