MCAT Foundations · Biology

DNA Structure and Replication

13 min read
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 5 sections
  1. In 30 seconds
  2. The college version
  3. Eli explains
  4. Study tools
  5. Sources & references

In 30 seconds

DNA is the universal information storage molecule of life — a double-stranded helix that encodes the complete blueprint for every protein a cell will ever make. The MCAT demands that you understand DNA at three levels: chemical (nucleotide composition, phosphodiester bonds, hydrogen bonding between bases), structural (antiparallel double helix, major and minor grooves, 5ʹ→3ʹ directionality), and functional (how the molecule is faithfully copied by the replisome). The central insight is that structure dictates function: the double helix suggests a mechanism for replication (each strand serves as a template), the antiparallel orientation explains why one strand is synthesized continuously and the other in fragments, and the chemistry of base pairing (A=T, G≡C) ensures fidelity. The MCAT connects DNA structure directly to replication enzymology — helicase, primase, DNA polymerases, ligase, topoisomerase — and asks you to predict what happens when components fail. Beyond replication, you must understand why linear chromosomes shorten with each round of replication (the end-replication problem) and how telomerase solves this in stem and germ cells. A recurring theme: everything in DNA biology flows from the 5ʹ→3ʹ polarity of synthesis and the complementary base-pairing rules.

The college version

The monomer unit of DNA is the deoxyribonucleotide, composed of three components: a nitrogenous base, a deoxyribose sugar (missing the 2ʹ-OH group found in RNA), and one to three phosphate groups. The nitrogenous bases fall into two families: purines (adenine and guanine, with fused double-ring structures) and pyrimidines (cytosine and thymine, with single six-membered rings). A nucleoside is the base-sugar combination without phosphate; a nucleotide is the base-sugar-phosphate unit. Nucleotides are linked by phosphodiester bonds — covalent linkages between the 3ʹ hydroxyl of one deoxyribose and the 5ʹ phosphate of the next — forming a sugar-phosphate backbone with a free 5ʹ end (phosphate) and a free 3ʹ end (hydroxyl). This 5ʹ→3ʹ polarity is the directional axis along which all DNA synthesis and reading occur. The MCAT often tests the distinction between nucleoside and nucleotide (nucleoside = no phosphate), the numbering of the sugar carbons (1ʹ attaches the base, 5ʹ attaches the phosphate), and the structural difference between ribose and deoxyribose at the 2ʹ position.

The Watson-Crick model of DNA (1953) describes a right-handed double helix with two antiparallel strands wound around a common axis. The sugar-phosphate backbones run along the exterior; the nitrogenous bases project inward and form specific hydrogen-bonded pairs: adenine pairs with thymine via two hydrogen bonds (A=T), and guanine pairs with cytosine via three hydrogen bonds (G≡C). This complementarity is the foundation of replication, transcription, and all DNA-based technologies. The helix has key dimensions: 2 nm in diameter, 10 base pairs per complete turn (3.4 nm pitch), with 0.34 nm between adjacent base pairs. The asymmetric stacking creates major and minor grooves — the major groove exposes more chemical information and is where most DNA-binding proteins (transcription factors, repair enzymes) read the sequence without unwinding the helix. Chargaff's rules state that in double-stranded DNA, %A = %T and %G = %C — a direct consequence of complementary base pairing. GC-rich DNA is more thermally stable (higher melting temperature, Tm) because of the extra hydrogen bond per G≡C pair. DNA is predominantly in the B-form under physiological conditions, but A-form (dehydrated, shorter and wider) and Z-form (left-handed, in alternating GC sequences) exist and may appear in MCAT passages.

The two strands of the DNA double helix run in opposite directions: one strand runs 5ʹ→3ʹ while its complement runs 3ʹ→5ʹ. This antiparallel arrangement is a direct consequence of the geometry of base pairing — the nucleotides must be oriented head-to-tail to satisfy the hydrogen-bonding distances. The antiparallel orientation has profound implications for replication: because all DNA polymerases synthesize exclusively in the 5ʹ→3ʹ direction (adding nucleotides to the 3ʹ-OH of the growing chain), one strand — the leading strand — can be copied continuously toward the replication fork, while the other — the lagging strand — must be copied discontinuously in short fragments as the fork opens. The MCAT tests antiparallel logic relentlessly: given a parental strand sequence and a replication fork direction, you must identify which strand is leading vs. lagging, predict the RNA primer locations, and determine the sequence of newly synthesized DNA. The rule is simple but the application trips students up: all new DNA synthesis is 5ʹ→3ʹ, so the template is always read 3ʹ→5ʹ.

DNA replicates semiconservatively: each daughter double helix contains one parental strand and one newly synthesized strand. This was demonstrated by the Meselson-Stahl experiment (1958), which used ¹⁵N (heavy) and ¹⁴N (light) isotopes of nitrogen to label DNA. E. coli grown in ¹⁵N medium were transferred to ¹⁴N medium and DNA was separated by equilibrium density gradient centrifugation. After one generation, all DNA banded at an intermediate density (ruling out conservative replication, which would have produced separate heavy and light bands). After two generations, equal amounts of intermediate and light DNA appeared (ruling out dispersive replication, which would have produced a single band of progressively lighter DNA). The experiment is a classic of MCAT experimental analysis: know the predicted outcomes for each model and what was actually observed.

The replication machinery (the replisome) is a multi-enzyme complex. Helicase unwinds the double helix by hydrolyzing ATP, breaking hydrogen bonds between base pairs and creating the replication fork. Single-strand binding proteins (SSBs) coat the separated strands to prevent reannealing and protect against nucleases. Topoisomerase (DNA gyrase in prokaryotes) relieves the positive supercoiling that builds ahead of the replication fork — it cuts one or both DNA strands, allows them to rotate, and reseals the break. Primase synthesizes short RNA primers (~10 nucleotides) that provide the free 3ʹ-OH required by DNA polymerase to begin synthesis. DNA polymerase III (prokaryotes) is the primary replicative enzyme — it extends primers, adding deoxynucleotides to the 3ʹ end with high processivity. DNA polymerase I removes RNA primers (via 5ʹ→3ʹ exonuclease activity) and fills the resulting gaps with DNA. DNA ligase seals the nicks between Okazaki fragments by catalyzing phosphodiester bond formation (using ATP or NAD⁺). In eukaryotes, the polymerase assignments differ: Pol α (with primase activity) initiates synthesis, Pol ε synthesizes the leading strand, and Pol δ synthesizes the lagging strand. All DNA polymerases have two critical properties: 5ʹ→3ʹ polymerase activity for synthesis and 3ʹ→5ʹ exonuclease activity for proofreading — they can detect and remove mismatched bases immediately after incorporation, reducing the error rate from ~10⁻⁵ to ~10⁻⁷ per base.

The asymmetry of the replication fork creates two mechanistically distinct strands. The leading strand is synthesized continuously: its template runs 3ʹ→5ʹ toward the fork, so a single RNA primer suffices, and DNA polymerase III extends it in the 5ʹ→3ʹ direction toward the advancing fork. The lagging strand is synthesized discontinuously: its template runs 5ʹ→3ʹ toward the fork, which forces synthesis away from the fork. As helicase exposes new template, primase lays down RNA primers at intervals (~1,000–2,000 nucleotides in prokaryotes, ~100–200 in eukaryotes), and DNA polymerase III synthesizes short DNA segments — Okazaki fragments — in the 5ʹ→3ʹ direction back toward the previously synthesized fragment. DNA polymerase I then removes the RNA primers and fills the gaps, and DNA ligase seals the nicks. The lagging strand loops out so that both polymerases can move together as a holoenzyme complex. A critical MCAT point: all DNA synthesis, on both strands, is 5ʹ→3ʹ — the difference is only in whether synthesis is continuous (leading) or fragmented (lagging).

Linear chromosomes face the end-replication problem: the lagging strand cannot replicate the extreme 3ʹ end, because removal of the final RNA primer leaves a gap that cannot be filled — there is no upstream 3ʹ-OH to extend from. This causes progressive shortening of chromosomes with each cell division, which ultimately triggers senescence or apoptosis (the Hayflick limit). Telomeres are repetitive, non-coding sequences (TTAGGG in humans, repeated thousands of times) at chromosome ends that buffer against this shortening. In stem cells, germ cells, and most cancer cells, telomerase — a ribonucleoprotein reverse transcriptase — extends telomeres by using its intrinsic RNA template (TERC) to add TTAGGG repeats to the 3ʹ overhang. The catalytic subunit (TERT) is the rate-limiting component; its expression determines whether a cell can maintain telomere length. DNA repair pathways preserve genomic integrity: proofreading (3ʹ→5ʹ exonuclease by DNA polymerase) catches errors during replication; mismatch repair (MMR) corrects errors that escape proofreading by recognizing distortions in the helix and targeting the newly synthesized strand (identified by transient nicks in prokaryotes, or by PCNA and strand discontinuities in eukaryotes); base excision repair (BER) removes damaged or inappropriate bases (e.g., uracil, oxidized bases) via glycosylases; nucleotide excision repair (NER) removes bulky helix-distorting lesions (e.g., thymine dimers from UV radiation) by excising a ~30-nucleotide patch. Defects in repair pathways are linked to cancer (Lynch syndrome = MMR defect; xeroderma pigmentosum = NER defect).

How it works

DNA Replication — Step by Step at the Fork

  1. Initiation: Initiator proteins bind the origin of replication (single oriC in prokaryotes; multiple origins in eukaryotes). Helicase is loaded onto DNA and unwinds the duplex bidirectionally, powered by ATP hydrolysis, creating two replication forks.
  1. Primer synthesis: Primase synthesizes a short RNA primer (~10 nt) complementary to the template, providing the essential free 3ʹ-OH. The leading strand requires one primer at the origin; the lagging strand requires a new primer for each Okazaki fragment.
  1. Elongation: DNA polymerase III (prokaryotes) or Pol ε/δ (eukaryotes) extends the primer by adding dNTPs complementary to the template. Nucleotide addition is driven by the hydrolysis of the incoming dNTP's β- and γ-phosphates (pyrophosphate release). Synthesis is processive: the sliding clamp (β-clamp in prokaryotes, PCNA in eukaryotes) tethers the polymerase to DNA.
  1. Lagging strand synthesis: The lagging strand template loops out so the polymerase moves with the replisome. Primase periodically synthesizes new primers. DNA polymerase III extends each primer until it reaches the previous Okazaki fragment.
  1. Primer removal and ligation: DNA polymerase I (prokaryotes) removes RNA primers via 5ʹ→3ʹ exonuclease activity and fills the gaps with DNA. DNA ligase then seals the remaining nicks, using ATP to adenylate the 5ʹ phosphate and drive phosphodiester bond formation with the 3ʹ-OH.
  1. Supercoil resolution: Topoisomerase (DNA gyrase) cuts the DNA ahead of the fork, allows controlled rotation to release torsional stress, and reseals. This prevents the fork from stalling due to overwound DNA.
  1. Termination: In prokaryotes, replication terminates at specific ter sequences where Tus proteins block helicase progression. In eukaryotes, forks from adjacent origins converge. Telomerase extends the 5ʹ ends of linear chromosomes in cells that express it.

Comparisons

  • Biochemistry: The phosphodiester bond — 5ʹ-phosphate to 3ʹ-OH, with pyrophosphate as the leaving group. The free energy of dNTP hydrolysis drives polymerization. Hydrogen bonding between bases (2 for A-T, 3 for G-C) explains thermal stability and influences PCR primer design.
  • Genetics/Molecular Biology: Semiconservative replication means mutations in one strand are perpetuated — the Meselson-Stahl experiment is a classic of molecular genetics. Telomerase is a reverse transcriptase, linking DNA replication to the central dogma (RNA → DNA).
  • Cell Biology: Telomere shortening and the Hayflick limit connect DNA replication to cellular senescence and the replicative lifespan of somatic cells. Cancer cells universally reactivate telomerase (or the ALT pathway) to achieve replicative immortality.
  • Pharmacology: Many chemotherapeutic agents target replication: topoisomerase inhibitors (etoposide, doxorubicin), nucleotide analogs (5-fluorouracil, cytarabine), and DNA-alkylating agents. These appear in passage-based questions.
  • Experimental Biology: PCR, Sanger sequencing, and next-generation sequencing all exploit 5ʹ→3ʹ synthesis, complementary base pairing, and the requirement for a 3ʹ-OH primer. Understanding DNA polymerase enzymology is essential for interpreting these techniques.

Common confusions

  • All synthesis is 5ʹ→3ʹ — always. Students incorrectly think the lagging strand is synthesized 3ʹ→5ʹ because it 'goes backward.' The lagging strand IS synthesized 5ʹ→3ʹ; it just does so in short, discontinuous segments. The template is always read 3ʹ→5ʹ.
  • DNA polymerase cannot initiate synthesis de novo. It requires a free 3ʹ-OH, provided by an RNA primer synthesized by primase. This is why the leading strand needs one primer (at the origin) but the lagging strand needs many. Missing this distinction leads to wrong answers about which enzymes are required first at the fork.
  • The Meselson-Stahl experiment tested three models, not two. Students often remember semiconservative vs. conservative but forget dispersive. After one generation: one intermediate band (rules out conservative). After two generations: one intermediate + one light band (rules out dispersive, which would produce only one band of intermediate-light density).
  • G≡C content affects DNA stability — know the numbers. GC base pairs have three hydrogen bonds vs. two for AT pairs. DNA with higher GC content has a higher melting temperature (Tm). Rank DNA fragments by Tm given their sequences — count the G+C percentage.
  • Telomerase is a reverse transcriptase, not a DNA polymerase. It carries its own RNA template and synthesizes DNA from RNA — the reverse of transcription. It extends the 3ʹ overhang of the parental strand, which then serves as a template for conventional lagging-strand synthesis by DNA polymerase. Active in stem cells, germ cells, and >85% of cancers, but not in most somatic cells.
  • Topoisomerase vs. helicase — they solve different problems. Helicase unwinds the duplex by breaking base-pair hydrogen bonds (ATP-powered). Topoisomerase relieves the torsional stress (supercoiling) caused by that unwinding — it breaks and reseals the sugar-phosphate backbone. DNA gyrase is a type II topoisomerase specific to prokaryotes; it is the target of fluoroquinolone antibiotics (e.g., ciprofloxacin).

Quick review

  • Nucleotide = nitrogenous base + deoxyribose + phosphate. Purines (A, G, double-ring) vs. pyrimidines (C, T, single-ring). Nucleoside = base + sugar (no phosphate). Phosphodiester bond: 3ʹ-OH to 5ʹ-phosphate.
  • Watson-Crick pairing: A=T (2 H-bonds), G≡C (3 H-bonds). B-form DNA: right-handed, 10 bp/turn, 2 nm diameter, major and minor grooves. Chargaff's rules: %A = %T, %G = %C.
  • Strands are antiparallel (5ʹ→3ʹ and 3ʹ→5ʹ). All DNA polymerases synthesize ONLY 5ʹ→3ʹ; template is read 3ʹ→5ʹ.
  • Semiconservative replication (Meselson-Stahl: ¹⁵N → intermediate, 2nd gen → intermediate + light). Each daughter duplex = one parental + one new strand.
  • Replisome enzymes: helicase (unwinds, ATP), SSBs (prevent reannealing), topoisomerase/gyrase (relieves supercoils), primase (RNA primers), DNA Pol III (main synthesis), DNA Pol I (removes primers, fills gaps), ligase (seals nicks, ATP).
  • Leading strand: continuous, one primer, toward fork. Lagging strand: discontinuous, Okazaki fragments, multiple primers. Synthesis always 5ʹ→3ʹ on both strands.
  • End-replication problem: lagging strand cannot replicate extreme 3ʹ end → telomere shortening. Telomerase = reverse transcriptase with built-in RNA template (TERC + TERT); extends TTAGGG repeats. Active in stem cells, germ cells, cancer cells.
  • DNA repair: proofreading (3ʹ→5ʹ exonuclease), mismatch repair (MMR, post-replication), base excision repair (BER, damaged bases), nucleotide excision repair (NER, bulky lesions like thymine dimers).
  • DNA Pol I activities: 5ʹ→3ʹ polymerase, 3ʹ→5ʹ exonuclease (proofreading), 5ʹ→3ʹ exonuclease (primer removal). Pol III is the main replicative polymerase in prokaryotes.
Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

DNA is like a twisted zipper with two matching sides. Each tooth of the zipper is a 'base' — A always matches with T, and G always matches with C. To copy DNA, the cell unzips the zipper down the middle using a machine called helicase, splitting it into two separate strands. Each strand now acts as a template: a copying machine called DNA polymerase moves along each strand, reading the bases and adding matching partners to build a new complementary strand. Because the copying machine only works in one direction (like a train that only goes forward on its track), one strand gets copied smoothly in one go — that's the 'leading strand.' The other strand has to be copied backward in short pieces that are stitched together later — that's the 'lagging strand.' The result: two identical zippers, each with one old side and one new side. At the very tips of chromosomes are protective caps called telomeres, like the plastic tips on shoelaces that keep them from fraying. Each time a cell divides, a bit of the telomere wears away — when it gets too short, the cell stops dividing. This is one reason we age at the cellular level.

Keep learning

Ready to build on this? Continue to the next lesson.

Study tools & related lessonsRelated

Sources & references

  1. OpenStax Biology 2e — Chapter 14: DNA Structure and Function — OpenStax / Rice University
  2. Molecular Biology of the Cell, 4th Edition — Chapter 5: DNA Replication, Repair, and Recombination — NCBI Bookshelf

This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.