Biology for AP Courses · Genes and Proteins

The Genetic Code

8 min read
Biological facts (codon assignments, sickle-cell base change) are standard commonly taught reference concepts; verify against current primary texts before high-stakes use.
Want it in plain words first? Jump to Eli explains — the same idea, no jargon.
On this page 9 sections
  1. In 30 seconds
  2. Why this matters
  3. The college version
  4. Eli explains
  5. Worked example
  6. Key takeaway
  7. Check yourself
  8. Study tools
  9. Sources & references

In 30 seconds

A gene is a stretch of DNA, but its product is a protein built from 20 amino acids, and DNA is built from only four nucleotides. How does a four-letter alphabet specify a twenty-letter language? The answer is the : nucleotides are read in groups of three (codons), each specifying one amino acid. With 4³ = 64 possible codons and only 20 amino acids, the code is degenerate (redundant) — most amino acids are specified by more than one .

The code was cracked in the early 1960s by Marshall Nirenberg (poly-U codes for phenylalanine), Gobind Khorana (defined repeating RNAs), and colleagues — honored with the 1968 Nobel Prize. The code is nearly universal across life, one of the strongest pieces of evidence for a common ancestor, and it is read in a fixed from a to a . This topic covers how the code works, why redundancy matters, and how its structure shapes the effects of mutations.

Why this matters

The genetic code is the operating manual for the central dogma (DNA → RNA → protein). Without it you cannot interpret a DNA sequence, predict a protein's amino-acid sequence, or understand what mutations do. It explains classic disease at the molecular level: sickle cell disease results from a single base change in the β-globin gene (commonly taught as GAG → GTG), swapping glutamic acid for valine at position 6 — one codon, one amino acid, one dramatic change in hemoglobin. Redundancy also explains "silent" mutations: a base change that leaves the amino acid unchanged usually causes no disease. In biotechnology, the code lets researchers translate sequenced genomes into proteins, design codon-optimized genes for mRNA vaccines and protein drugs, and interpret genetic-testing variants. For the AP exam, codon-reading and mutation-effect questions are guaranteed territory.

The college version

Core Concepts

Triplets, non-overlapping, and a fixed reading frame

Three nucleotides = one codon = one amino acid. Evidence came from frameshift mutations: Crick and Brenner showed that adding or deleting 1 or 2 bases destroys gene function, but adding or deleting 3 often restores it — the message is read in groups of three. The code is non-overlapping: each nucleotide belongs to exactly one codon, so a sequence can be read in three possible reading frames; the correct one is set by the start codon. A single-base insertion or deletion shifts the frame and changes every downstream codon — a — usually destroying the protein.

How the code was cracked

  • Nirenberg's poly-U experiment (1961): poly(U) (UUUUUU…) in a cell-free translation system produced a protein made only of phenylalanine — first codon cracked: UUU = Phe.
  • Khorana's repeating polymers: RNA like UCUCUCUC… produced alternating Ser-Leu-Ser-Leu, proving non-overlapping triplets (UCU and CUC).
  • Triplet binding assays: Nirenberg and Leder showed specific RNA triplets make ribosomes bind specific aminoacyl-tRNAs — all 64 codons were quickly assigned.

Start and stop signals

AUG is the start codon — it codes for methionine and sets the reading frame (bacteria use a special formyl-methionine at the start). Three codons — UAA, UAG, UGA — are stop codons: they code for no amino acid and are recognized by release factors that detach the finished protein. Because translation runs from AUG in multiples of three to a stop, protein length is set by the distance between them.

Degeneracy and wobble

Most amino acids are encoded by 2–6 codons, usually sharing the first two bases and differing only in the third. Crick's hypothesis explains why this works: pairing between the tRNA and the mRNA codon is relaxed at the codon's third position, so one tRNA can often recognize several codons. The code is degenerate but not ambiguous: each codon always specifies one amino acid — redundancy, not vagueness. Only Met (AUG) and Trp (UGG) have single codons.

Near-universality and its exceptions

With only minor variations, all organisms use the same code — strong evidence for common ancestry. Known exceptions are mostly mitochondrial genomes (e.g., UGA codes for tryptophan instead of stop in vertebrate mitochondria) and a few organisms such as Mycoplasma. The exceptions are small, but they show the code is not quite frozen — close enough, however, that a human gene can be read and even expressed by bacteria, the basis of recombinant DNA technology.

Common Confusions

Do Not ConfuseWithThe Difference
CodonAnticodonCodon is on mRNA; anticodon is on tRNA and base-pairs with the codon
Degenerate codeAmbiguous codeDegenerate = redundant (several codons → one amino acid); the code is never ambiguous — each codon specifies exactly one amino acid
DNA codeRNA codeSame code, but DNA uses T where RNA uses U; codons are written in mRNA (with U)
64 codons64 amino acidsThere are 64 codons but only 20 amino acids; 3 codons are stops
AUG (start)Stop codonsAUG codes for methionine and starts translation; UAA/UAG/UGA code for nothing and stop it
Frameshift mutationPoint (substitution) mutationFrameshifts change every codon after the indel; substitutions change only one codon
Reading frameReading directionFrame = which nucleotides group into codons; direction = 5′→3′ — both matter, but they are different concepts
Eli, the EliExplains learning guide

Eli explains

The same idea, in plain words

Explain it like I’m 10

The genetic code is like a secret language with only 4 letters that writes sentences using 3-letter words. Your cells read the DNA story three letters at a time, and each 3-letter word names one LEGO brick (an amino acid). There are 64 possible words but only 20 bricks, so most bricks have several words that mean the same thing. One word, AUG, means "start reading here," and three words mean "stop." Every living thing on Earth reads the same dictionary — a clue that all life shares one ancestor.

Worked example

Think of the code as a sentence of three-letter words: THE CAT ATE THE RAT AND RAN. Now delete the first letter: HEC ATA TET HER ATA NDR AN — every word is nonsense after the deletion point. That is a frameshift mutation: one deleted base changes every downstream codon, and the resulting protein is garbage. Now change just one letter in the middle: THE CAT ATE THE MAT AND RAN — one word changed, the rest untouched. That is a missense (point) mutation: a single amino acid substitution. Sickle cell disease is the classic real example: in the β-globin gene, the commonly taught change GAG → GTG converts the sixth amino acid from glutamic acid (hydrophilic) to valine (hydrophobic); in deoxygenated blood, mutant hemoglobin molecules stick together and deform red blood cells into sickles — one codon, one amino acid, and a whole disease. Finally, GAG → GAA still specifies glutamic acid: a silent mutation, thanks to degeneracy.

Key takeaways

  • Codons are triplets: 3 nucleotides = 1 amino acid; 4³ = 64 codons; 61 sense + 3 stop (UAA, UAG, UGA).
  • AUG = start = methionine; it sets the reading frame.
  • The code is degenerate (redundant): most amino acids have 2–6 codons; only Met and Trp have one.
  • The code is non-overlapping, read 5′→3′ in a fixed frame; indels not in multiples of 3 cause frameshifts that scramble everything downstream.
  • Degenerate ≠ ambiguous: each codon specifies exactly one amino acid.
  • Wobble (relaxed pairing at the codon's 3rd base) lets fewer than 61 tRNAs read all sense codons.
  • The code is nearly universal; minor exceptions occur in mitochondria and a few prokaryotes.
  • Key experiments: Nirenberg's poly-U (UUU = Phe) and Khorana's repeating polymers proved triplet, non-overlapping reading.
  • Sickle cell disease: the commonly taught single-base change GAG → GTG in β-globin changes Glu → Val at position 6 — one codon, altered protein behavior.

Check yourself

6 review questions from the chapter. Try each one, then open the answer.

  1. Why must the code use at least three nucleotides per amino acid?

    Show answer

    Four nucleotides give 4¹ = 4 one-base words and 4² = 16 two-base words — not enough for 20 amino acids; triplets give 4³ = 64, more than enough.

  2. How many codons specify amino acids, and how many are stop codons? Name them.

    Show answer

    61 codons specify amino acids; 3 are stop codons: UAA, UAG, UGA. AUG doubles as the start codon (methionine).

  3. What did Nirenberg's poly-U experiment demonstrate?

    Show answer

    That UUU (poly-U RNA) directs synthesis of a protein made only of phenylalanine — the first codon was cracked, proving RNA sequence determines amino acid sequence.

  4. Why is a base-pair deletion usually more damaging than a base-pair substitution?

    Show answer

    A substitution changes one codon, leaving the rest of the protein intact. A deletion shifts the reading frame, so every downstream codon is regrouped and usually all subsequent amino acids change — typically destroying protein function.

  5. What is wobble, and what does it allow?

    Show answer

    Wobble is the relaxed base-pairing between the third base of the codon and the first base of the anticodon. It lets one tRNA recognize several codons differing only in the third position, so fewer than 61 tRNAs are needed.

  6. Why is the near-universality of the genetic code considered evidence for common ancestry?

    Show answer

    Independent invention would almost certainly produce different codes; essentially all life using the same assignments is best explained by descent from a common ancestor.

Keep learning

Ready to build on this? Continue to the next lesson.

Study tools & related lessonsKey vocabulary · Related

Key vocabulary

codon
A three-nucleotide unit of mRNA specifying one amino acid (or stop)
genetic code
The complete mapping of all 64 codons to amino acids and stops
reading frame
The grouping of nucleotides into codons, set by the start codon
start codon
AUG; initiates translation and codes for methionine
stop codon
UAA, UAG, or UGA; codes for no amino acid
degeneracy (redundancy)
Most amino acids are specified by more than one codon
wobble
Relaxed base-pairing between the codon's 3rd base and anticodon's 1st base
anticodon
The three nucleotides on a tRNA that pair with the codon
frameshift mutation
Insertion/deletion of a number of bases not divisible by 3

Sources & references

  1. openstax.org — Biology Ap Courses

This lesson was adapted from the open educational references above; their licenses and attributions are preserved. See Copyright & Licensing.

Educational content only. It is not medical, legal or professional advice. Found an error? Tell us.