Biology for AP Courses · DNA Structure and Function

DNA Structure and Sequencing

8 min read
Helix dimensions, human genome size, gene-count estimates, and genome-sequencing timelines are commonly taught reference concepts from introductory biology; verify against current primary sources (e.g., NCBI, NIH) before formal citation.
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 9 sections
  1. In 30 seconds
  2. Why this matters
  3. The college version
  4. Eli explains
  5. Worked example
  6. Key takeaway
  7. Check yourself
  8. Study tools
  9. Sources & references

In 30 seconds

Deoxyribonucleic acid (DNA) is a long polymer built from repeating units called nucleotides. Each has three parts: a phosphate group, a five-carbon sugar (deoxyribose), and one of four nitrogenous bases — adenine (A), thymine (T), guanine (G), and cytosine (C). Nucleotides link into a strand with a built-in direction, and two strands wind around each other to form the double helix, held together by hydrogen bonds between paired bases (A with T, G with C) and running in directions. This structure explains how DNA is copied, how it is read into RNA, and how mutations arise.

is the process of determining the exact order of bases along a DNA molecule. Knowing a sequence lets scientists identify genes, compare organisms, trace evolutionary relationships, and match crime-scene samples to suspects. Sequencing technology has driven a revolution: the first human genome took roughly a decade to complete; modern instruments can sequence a human genome in about a day (commonly taught figures that change rapidly — verify against current sources). Structure and sequencing are two halves of the same story.

Why this matters

  • Structure explains function: Base pairing (A–T, G–C) is the mechanism behind replication, transcription, and mutation — the topics this chapter and the next build on.
  • Forensics and identity: Sequence differences between people are the basis of DNA profiling used in criminal justice and paternity testing.
  • Medicine: Sequencing identifies disease-causing mutations, guides cancer treatment choices, and underlies prenatal and newborn screening.
  • Evolution: Comparing DNA sequences across species is now the standard way to build evolutionary trees and track pathogens.
  • Exams: Expect questions on nucleotide anatomy, base-pairing rules, DNA vs. RNA differences, directionality (5′ and 3′), and the logic of Sanger sequencing.

The college version

Core Concepts

Nucleotides: the building blocks

Each nucleotide consists of a phosphate group (which gives DNA its negative charge), a five-carbon sugar (deoxyribose; the carbons are numbered 1′–5′, and "deoxy" means it lacks an –OH at the 2′ carbon), and a nitrogenous base attached to the 1′ carbon. Nucleotides join into a strand by phosphodiester bonds between the 5′ phosphate of one nucleotide and the 3′ –OH of the next. The chain therefore has a 5′ end (free phosphate) and a 3′ end (free hydroxyl). This directionality is essential — polymerases build DNA only in the 5′ → 3′ direction, a rule that drives the entire replication story.

Purines and pyrimidines

The four bases come in two shapes. Purines (adenine, guanine) have two fused rings; pyrimidines (thymine, cytosine) have one. In the double helix a purine always pairs with a pyrimidine, which keeps the two sugar-phosphate backbones a constant distance apart. Pairing is specific and held by hydrogen bonds: A pairs with T (two hydrogen bonds) and G pairs with C (three). Because G–C pairs have one more hydrogen bond, G–C-rich DNA is slightly harder to separate — a fact exploited in PCR and melting-temperature calculations.

The double helix

Watson and Crick's 1953 model (built on Franklin's X-ray data and Chargaff's rules) described DNA as two antiparallel strands — one running 5′ → 3′, the other 3′ → 5′ — wound around a common axis. The sugar-phosphate backbones form the outside; the paired bases stack inside like the rungs of a twisted ladder. Commonly taught dimensions: the helix is about 2 nm in diameter, bases are stacked about 0.34 nm apart, and there are about 10 base pairs per turn. The strands are — the sequence of one determines the sequence of the other — which is the basis of copying and hybridization.

DNA versus RNA

RNA differs from DNA in three ways: its sugar is ribose (with an –OH at the 2′ carbon), it uses uracil (U) in place of thymine (U pairs with A), and it is usually single-stranded (though it folds into complex shapes like tRNA cloverleaves and ribosomes). RNA's extra hydroxyl makes it less stable than DNA — a trade-off that suits its many short-lived roles.

Packaging the genome

A human cell's DNA is roughly 2 meters long if fully extended (a commonly taught figure) yet fits into a nucleus about 6 µm across — a compression problem solved by wrapping DNA around histone proteins to form nucleosomes, which coil into chromatin and condense into chromosomes during division. The human genome contains about 3 billion base pairs per haploid set, with roughly 20,000–25,000 protein-coding genes (commonly taught estimates; the exact count is still being refined). Additional twisting (supercoiling) further compacts DNA and is managed by enzymes called topoisomerases.

Sequencing DNA

Sanger (chain-termination) sequencing reads a DNA molecule by copying it in the presence of dideoxynucleotides (ddNTPs). A ddNTP can be added to a growing strand but, lacking the 3′ –OH, cannot be extended — it terminates the chain. With fluorescently labeled ddNTPs, each base produces a chain of a characteristic length, and separating the fragments by size while reading the colors reveals the sequence. Next-generation sequencing (NGS) modernized the idea with massively parallel sequencing-by-synthesis: millions of short fragments are anchored to a surface and read base-by-base as fluorescent signals, generating billions of reads per run that computers assemble into genomes. The same technology powers clinical panels, tumor profiling, and metagenomics.

Common Confusions

Do Not ConfuseWithDifference
5′ and 3′ endsInterchangeable endsThe 5′ end has a free phosphate, the 3′ end a free hydroxyl; synthesis always extends the 3′ end.
A–T vs. G–C pairing strengthAll pairs equally easy to separateG–C pairs share three hydrogen bonds, A–T only two, so G–C-rich DNA denatures at higher temperatures.
DNARNARNA uses ribose and uracil and is usually single-stranded; DNA uses deoxyribose and thymine and is double-stranded.
GenomeGeneA genome is the entire DNA content (≈3 billion bp in humans); a gene is one functional unit within it.
SequencingPCR (amplification)Sequencing reads the order of bases; PCR copies a specific segment many times (often a step before sequencing).
Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

DNA is a recipe book written in a four-letter alphabet — A, T, G, and C — and the letters always pair up: A with T, and G with C. Two strands of letters stick together like a zipper, and the pairs stack up like the steps of a spiral staircase. To "sequence" DNA means to read the order of the letters, one by one, like reading a book from front to back — and computers help scientists read millions of letters at once.

Worked example

Suppose you are given a short DNA strand and asked to determine its sequence: 5′ – G A T T A C – 3′. The reasoning mirrors what sequencing instruments do:

  1. Find the complementary strand. The partner must be 3′ – C T A A T G – 5′, because G pairs with C and A pairs with T.
  2. Think about direction. A polymerase copying the top strand builds the bottom strand 5′ → 3′, adding C, then T, then A, then A, then T, then G.
  3. Predict Sanger output. The ddNTP for each base terminates chains at the positions of that base; the set of fragment lengths, read as a ladder of colors, spells out the sequence 5′ → 3′.
  4. Scale up. A real genome is the same logic repeated billions of times, with computers assembling overlapping short reads into whole chromosomes — and comparing the result against reference genomes to find variants that matter for health or identity.

Key takeaways

  • A nucleotide = phosphate + deoxyribose + nitrogenous base; A, G are purines (two rings); C, T are pyrimidines (one ring).
  • Strands are built 5′ → 3′; the double helix is antiparallel and complementary.
  • Base pairing: A–T (2 H-bonds), G–C (3 H-bonds); purine always pairs with pyrimidine.
  • Commonly taught dimensions: helix ≈ 2 nm diameter, 0.34 nm per stacked base pair, ~10 bp per turn.
  • DNA vs. RNA: deoxyribose vs. ribose, T vs. U, double-stranded vs. usually single-stranded.
  • Human genome ≈ 3 billion bp, ~20,000–25,000 protein-coding genes (commonly taught estimates; verify against current sources).
  • Sanger sequencing uses ddNTPs that stop chain growth; NGS runs the same idea massively in parallel.

Check yourself

5 review questions from the chapter. Try each one, then open the answer.

  1. Name the three parts of a nucleotide and the four bases of DNA.

    Show answer

    Phosphate group, deoxyribose sugar, and a nitrogenous base. The bases are adenine (A), thymine (T), guanine (G), and cytosine (C).

  2. Why must a purine always pair with a pyrimidine in the double helix?

    Show answer

    A purine (two rings) pairing with a pyrimidine (one ring) keeps the sugar-phosphate backbones a constant distance apart, giving the helix its uniform ~2 nm width. Purine–purine pairs would bulge; pyrimidine–pyrimidine pairs would pinch.

  3. What makes the two strands of DNA "antiparallel," and why does directionality matter?

    Show answer

    One strand runs 5′ → 3′ while the other runs 3′ → 5′. Directionality matters because polymerases synthesize DNA only in the 5′ → 3′ direction, which shapes how replication and transcription work.

  4. List three differences between DNA and RNA.

    Show answer

    RNA uses ribose instead of deoxyribose, uracil instead of thymine, and is usually single-stranded rather than double-stranded.

  5. How does a dideoxynucleotide make Sanger sequencing work?

    Show answer

    A ddNTP lacks the 3′ hydroxyl, so once it is added to a growing strand no further nucleotides can attach — the chain terminates. The lengths of terminated fragments reveal the position of each base.

Keep learning

Ready to build on this? Continue to the next lesson.

Study tools & related lessonsKey vocabulary · Related

Key vocabulary

Nucleotide
Phosphate + 5-carbon sugar + nitrogenous base
Purine / Pyrimidine
Two-ring bases (A, G) / one-ring bases (C, T)
Phosphodiester bond
Link between the 5′ phosphate and 3′ hydroxyl of adjacent nucleotides
5′ / 3′ ends
The two directional ends of a nucleotide chain
Antiparallel
The two strands run in opposite directions
Complementary
One strand's sequence determines the other's (A–T, G–C)
Dideoxynucleotide (ddNTP)
Nucleotide that terminates chain growth
Sequencing
Reading the order of bases in a DNA molecule
Nucleosome
DNA wrapped around histone proteins

Sources & references

  1. openstax.org — Biology Ap Courses

This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.