MCAT Foundations · Biology

Gene Expression: Transcription and RNA Processing

10 min read
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 5 sections
  1. In 30 seconds
  2. The college version
  3. Eli explains
  4. Study tools
  5. Sources & references

In 30 seconds

Gene expression is the process by which information encoded in DNA is converted into functional gene products — proteins or noncoding RNAs. The first major step, transcription, uses DNA as a template to synthesize a complementary RNA strand. In eukaryotes, this primary transcript (pre-mRNA) then undergoes extensive RNA processing — capping, polyadenylation, and splicing — before it becomes a mature mRNA ready for translation. Prokaryotes transcribe and translate simultaneously in the cytoplasm without these processing steps. Understanding transcription and RNA processing is essential for the MCAT: it ties together molecular biology, genetic regulation, and the molecular basis of disease (e.g., splicing defects in beta-thalassemia). The central dogma — DNA → RNA → protein — hinges on the fidelity and regulation of these steps.


The college version

1. Transcription Machinery

Transcription is catalyzed by RNA polymerases, which synthesize RNA in the 5′→3′ direction using a DNA template strand (read 3′→5′). Unlike DNA polymerases, RNA polymerases do not require a primer. In prokaryotes, a single RNA polymerase core enzyme associates with a sigma (σ) factor to form the holoenzyme, which recognizes promoter sequences. In eukaryotes, three distinct RNA polymerases exist: RNA Pol I (rRNA), RNA Pol II (mRNA, snRNA, miRNA), and RNA Pol III (tRNA, 5S rRNA). RNA Pol II is the focus for protein-coding gene transcription. Transcription occurs in three phases: initiation (polymerase binds promoter and unwinds DNA), elongation (polymerase moves along the template, adding ribonucleotides complementary to the template strand), and termination (polymerase dissociates at terminator sequences). In prokaryotes, termination can be Rho-dependent or Rho-independent (hairpin loop formation). In eukaryotes, termination is coupled to polyadenylation signals (AAUAAA).

2. Promoters, Enhancers, and Transcription Factors

Promoters are DNA sequences immediately upstream of the transcription start site that serve as binding platforms for RNA polymerase and general transcription factors (GTFs). The core promoter contains elements such as the TATA box (TATAAAA, ~−25 to −30 bp upstream), the initiator (Inr), and downstream promoter elements (DPE). General transcription factors (TFIIA, TFIIB, TFIID, TFIIE, TFIIF, TFIIH) assemble with RNA Pol II at the promoter to form the pre-initiation complex (PIC). TFIID contains the TATA-binding protein (TBP) subunit that recognizes the TATA box. TFIIH has helicase activity that unwinds DNA and kinase activity that phosphorylates the C-terminal domain (CTD) of RNA Pol II, triggering promoter escape and elongation.

Enhancers are distal regulatory DNA elements that can be located thousands of base pairs upstream, downstream, or even within introns. They bind activator proteins (specific transcription factors) that loop DNA to contact the PIC via the Mediator complex, dramatically increasing transcription rates. Silencers are analogous elements that bind repressor proteins to decrease transcription. Enhancer-promoter looping explains how distant regulatory elements influence gene expression — a key concept for understanding tissue-specific gene regulation.

3. RNA Polymerase

RNA Polymerase II is the central enzyme for mRNA synthesis in eukaryotes. It is a large, multi-subunit complex (~12 subunits in yeast). Its C-terminal domain (CTD), composed of heptapeptide repeats (Tyr-Ser-Pro-Thr-Ser-Pro-Ser), serves as a dynamic scaffold. The phosphorylation state of the CTD dictates which processing factors are recruited: unphosphorylated CTD recruits capping enzymes during initiation; Ser5-phosphorylated CTD associates with capping machinery; Ser2-phosphorylated CTD recruits splicing and polyadenylation factors during elongation. This CTD code couples transcription to RNA processing — ensuring that capping, splicing, and polyadenylation occur co-transcriptionally.

Key distinction: RNA Pol II is inhibited by α-amanitin (a toxin from the death cap mushroom, Amanita phalloides), a classic MCAT fact. RNA Pol I and III are resistant at low concentrations.

4. mRNA Processing

Eukaryotic pre-mRNA undergoes three major processing events before export to the cytoplasm:

  • 5′ Capping: A 7-methylguanosine (m⁷G) cap is added to the 5′ end via a 5′-to-5′ triphosphate linkage. The cap protects mRNA from 5′ exonucleases, aids nuclear export, and is recognized by eIF4E for translation initiation.
  • 3′ Polyadenylation: The pre-mRNA is cleaved ~10–30 nucleotides downstream of the AAUAAA polyadenylation signal, and poly(A) polymerase (PAP) adds ~200 adenine residues (the poly(A) tail). The tail is bound by poly(A)-binding protein (PABP) and protects the 3′ end from degradation while facilitating translation via the closed-loop model (PABP interacts with eIF4G → eIF4E → 5′ cap).
  • Splicing: Introns are removed and exons are ligated together (see Subtopics 5–6).

These three modifications are unique to eukaryotes and are prerequisites for mRNA export through the nuclear pore complex. Prokaryotic mRNA is not processed and is translated while still being transcribed.

5. Introns, Exons, and Splicing

Exons are the coding (and untranslated) sequences retained in mature mRNA. Introns are intervening sequences removed during splicing. Introns begin with a GU dinucleotide (5′ splice site or donor site) and end with an AG dinucleotide (3′ splice site or acceptor site) — the GU-AG rule. Upstream of the 3′ splice site is a branch point adenine within a consensus sequence, followed by a polypyrimidine tract.

Splicing is catalyzed by the spliceosome, a large ribonucleoprotein (RNP) complex composed of five small nuclear RNAs (U1, U2, U4, U5, U6 snRNAs) and hundreds of proteins. These snRNAs associate with proteins to form snRNPs (small nuclear ribonucleoproteins). The mechanism proceeds via two transesterification reactions:

  1. The branch point adenine 2′-OH attacks the 5′ splice site phosphate, forming a lariat intermediate.
  2. The free 3′-OH of the upstream exon attacks the 3′ splice site, ligating the exons and releasing the intron lariat.

The lariat intron is subsequently debranched and degraded.

6. Alternative Splicing

Alternative splicing allows a single gene to produce multiple mRNA isoforms — and therefore multiple protein variants — by differentially including or excluding exons from the final transcript. This is a major source of proteomic diversity: the human genome (~20,000 protein-coding genes) produces an estimated >100,000 proteins primarily through alternative splicing.

Key patterns include:

  • Exon skipping (cassette exon): An exon is either included or excluded.
  • Mutually exclusive exons: Only one of two exons is retained.
  • Alternative 5′ or 3′ splice sites: Splicing occurs at different donor or acceptor sites within an exon, changing its boundaries.
  • Intron retention: An intron is retained in mature mRNA (common in plants, less so in vertebrates).

Splicing is regulated by splicing enhancers and silencers (exonic or intronic: ESE, ISE, ESS, ISS) that bind SR proteins (serine/arginine-rich activators) and hnRNPs (heterogeneous nuclear ribonucleoproteins, typically repressors). Tissue-specific expression of splicing factors underlies differential isoform production — a classic example is the Dscam gene in Drosophila, which can generate >38,000 isoforms through alternative splicing.

7. Noncoding RNAs

Not all transcribed RNAs encode proteins. Noncoding RNAs (ncRNAs) are functional RNA molecules transcribed from DNA but not translated. Major classes relevant to the MCAT:

  • Ribosomal RNA (rRNA): Structural and catalytic component of ribosomes. Transcribed by RNA Pol I (28S, 18S, 5.8S) and RNA Pol III (5S).
  • Transfer RNA (tRNA): Adaptor molecules that decode mRNA codons into amino acids. Transcribed by RNA Pol III.
  • Small nuclear RNA (snRNA): Core components of the spliceosome (U1, U2, U4, U5, U6). Transcribed by RNA Pol II or III.
  • MicroRNA (miRNA): ~22-nucleotide RNAs that regulate gene expression post-transcriptionally by base-pairing with target mRNAs, leading to translational repression or degradation via the RNA-induced silencing complex (RISC).
  • Small interfering RNA (siRNA): Similar to miRNA, derived from exogenous double-stranded RNA; triggers RNA interference (RNAi).
  • Long noncoding RNA (lncRNA): >200 nucleotides; diverse roles including X-chromosome inactivation (Xist), chromatin remodeling, and transcriptional regulation.
  • Small nucleolar RNA (snoRNA): Guide chemical modifications (methylation, pseudouridylation) of rRNA and other RNAs in the nucleolus.

The discovery that most of the human genome is transcribed into ncRNAs — not protein-coding mRNAs — fundamentally reshaped our understanding of gene expression and genome function.


How it works

The transcriptional cycle can be understood as a coordinated sequence:

  1. Chromatin remodeling: Before transcription can begin, chromatin must be opened. Histone acetyltransferases (HATs) acetylate lysine residues on histones, neutralizing their positive charge and loosening DNA-histone interactions. ATP-dependent remodeling complexes (SWI/SNF) reposition nucleosomes to expose promoter regions.
  1. Assembly of the PIC: General transcription factors (TFIID through TFIIH) assemble at the core promoter in an ordered pathway. TFIID (via TBP) binds and bends the TATA box. TFIIH unwinds ~10–15 bp of DNA at the transcription start site using its helicase subunits, forming the open complex.
  1. Promoter escape: TFIIH phosphorylates Ser5 of the CTD. RNA Pol II clears the promoter, and the capping enzyme is recruited to modify the nascent 5′ end when the transcript is ~20–30 nucleotides long.
  1. Elongation: As RNA Pol II moves along the template, the elongation rate is ~20–50 nucleotides per second. Nucleosomes are disassembled ahead of the polymerase and reassembled behind it. The CTD phosphorylation pattern shifts from Ser5 to Ser2, recruiting splicing factors and polyadenylation factors.
  1. Co-transcriptional processing: Splicing and polyadenylation occur while the transcript is still tethered to RNA Pol II via the CTD. The polyadenylation signal (AAUAAA) triggers cleavage and release of the transcript.
  1. Export: Mature mRNA, bound by exon junction complexes (EJCs) and the cap-binding complex, is exported through the nuclear pore complex. The EJC marks successfully spliced transcripts; unspliced or improperly processed mRNAs are retained and degraded by the nuclear exosome.

How it works

The transcriptional cycle can be understood as a coordinated sequence:

  1. Chromatin remodeling: Before transcription can begin, chromatin must be opened. Histone acetyltransferases (HATs) acetylate lysine residues on histones, neutralizing their positive charge and loosening DNA-histone interactions. ATP-dependent remodeling complexes (SWI/SNF) reposition nucleosomes to expose promoter regions.
  1. Assembly of the PIC: General transcription factors (TFIID through TFIIH) assemble at the core promoter in an ordered pathway. TFIID (via TBP) binds and bends the TATA box. TFIIH unwinds ~10–15 bp of DNA at the transcription start site using its helicase subunits, forming the open complex.
  1. Promoter escape: TFIIH phosphorylates Ser5 of the CTD. RNA Pol II clears the promoter, and the capping enzyme is recruited to modify the nascent 5′ end when the transcript is ~20–30 nucleotides long.
  1. Elongation: As RNA Pol II moves along the template, the elongation rate is ~20–50 nucleotides per second. Nucleosomes are disassembled ahead of the polymerase and reassembled behind it. The CTD phosphorylation pattern shifts from Ser5 to Ser2, recruiting splicing factors and polyadenylation factors.
  1. Co-transcriptional processing: Splicing and polyadenylation occur while the transcript is still tethered to RNA Pol II via the CTD. The polyadenylation signal (AAUAAA) triggers cleavage and release of the transcript.
  1. Export: Mature mRNA, bound by exon junction complexes (EJCs) and the cap-binding complex, is exported through the nuclear pore complex. The EJC marks successfully spliced transcripts; unspliced or improperly processed mRNAs are retained and degraded by the nuclear exosome.

Comparisons

  • Biochemistry: Understanding the CTD code bridges molecular biology and enzymology. Phosphorylation, protein-protein interactions, and allosteric regulation — all core biochemistry concepts — converge at the CTD.
  • Genetics: Splicing mutations are a common cause of genetic disease. For example, a point mutation creating a cryptic splice site in the β-globin gene causes β-thalassemia. Intronic mutations disrupting the branch point or polypyrimidine tract abolish normal splicing.
  • Evolution: Alternative splicing explains how organismal complexity can increase without proportional increases in gene number. Comparing splicing patterns across species reveals evolutionary conservation of regulatory networks.
  • Pharmacology: Several antibiotics target transcription. Rifampin binds bacterial RNA polymerase and blocks initiation (used for tuberculosis). α-Amanitin inhibits eukaryotic RNA Pol II — a classic laboratory tool.
  • Cell Biology: The nuclear pore complex selectively exports mature mRNPs. This compartmentalization (transcription in nucleus, translation in cytoplasm) is a defining eukaryotic feature.
Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

Imagine your DNA is a giant cookbook locked in a vault (the nucleus). You can't take the whole cookbook out — it's too precious. So you make a photocopy of just the recipe you need. That photocopy is transcription — RNA polymerase copies a gene into messenger RNA.

But the photocopy comes out messy: it has extra pages (introns) mixed in with the real recipe (exons). Before you can use it in the kitchen (the cytoplasm), you need to edit the copy: add a protective cover (5′ cap), staple the end (poly-A tail), and cut out the junk pages then tape the good ones together (splicing). That's RNA processing.

Even cooler, one recipe can make different dishes depending on which pages you keep — that's alternative splicing. The same gene in a brain cell and a muscle cell can be spliced differently to make proteins suited to each job. And some RNAs aren't recipes at all — they're tools (noncoding RNAs) that help the kitchen run, like timers (miRNA), whisks (rRNA), and order slips (tRNA).


Keep learning

Ready to build on this? Continue to the next lesson.

Study tools & related lessonsRelated

Sources & references

  1. OpenStax Biology 2e — Chapter 15: Genes and Proteins — OpenStax (Rice University)
  2. NCBI Bookshelf: Molecular Biology of the Cell, 4th Edition — Chapter 6: How Cells Read the Genome: From DNA to Protein — National Center for Biotechnology Information (NCBI / NIH)

This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.