Skip to content

Botany · Ch 7 — Molecular Basis of Inheritance

Genetic Code

7.6

Genetic Code

The genetic code is the set of rules by which information encoded in genetic material (DNA or RNA) is translated into proteins. It is the language that links the sequence of nucleotides in mRNA to the sequence of amino acids in a polypeptide chain.

The code is written in triplets — every three consecutive nucleotides in mRNA form a codon, and each codon specifies one particular amino acid. Since there are four different nucleotides (A, U, G, C), the total number of possible triplets is four cubed, or 64. Out of these 64 codons, 61 code for amino acids, and the remaining three are stop codons (UAA, UAG, UGA) that signal the termination of protein synthesis. One codon, AUG, has a dual role: it codes for the amino acid methionine and also serves as the start codon (initiation codon) for translation.

The genetic code is universal — with very minor exceptions, the same codon specifies the same amino acid in all organisms, from bacteria to humans. This universality strongly supports the idea that all life shares a common evolutionary origin.

The code is degenerate — meaning that most amino acids are specified by more than one codon. For example, the amino acid leucine is coded by six different codons (UUA, UUG, CUU, CUC, CUA, CUG). Only two amino acids, methionine and tryptophan, are coded by a single codon each. Degeneracy helps minimise the harmful effects of mutations: a change in the third nucleotide of a codon often still codes for the same amino acid.

The code is non-overlapping and commaless. It is read sequentially from a fixed starting point, without any gaps or punctuation between codons. Each nucleotide belongs to exactly one codon, and the reading frame is set by the start codon.

The code is unambiguous — each codon specifies only one amino acid (or a stop signal). There is no ambiguity: a given codon never codes for more than one amino acid.

Note

The genetic code was deciphered by Marshall Nirenberg, Heinrich Matthaei, and others in the 1960s. Nirenberg received the Nobel Prize in 1968 for this work. They used cell-free systems and synthetic RNA molecules (like poly-U, which codes for phenylalanine) to crack the code.

The following table summarises the standard genetic code. The first base of the codon is on the left, the second base at the top, and the third base on the right.

First base (5')Second baseThird base (3')
UUUU – Phe, UUC – Phe, UUA – Leu, UUG – LeuU
UUCU – Ser, UCC – Ser, UCA – Ser, UCG – SerC
UUAU – Tyr, UAC – Tyr, UAA – Stop, UAG – StopA
UUGU – Cys, UGC – Cys, UGA – Stop, UGG – TrpG
CCUU – Leu, CUC – Leu, CUA – Leu, CUG – LeuU
CCCU – Pro, CCC – Pro, CCA – Pro, CCG – ProC
CCAU – His, CAC – His, CAA – Gln, CAG – GlnA
CCGU – Arg, CGC – Arg, CGA – Arg, CGG – ArgG
AAUU – Ile, AUC – Ile, AUA – Ile, AUG – Met (Start)U
AACU – Thr, ACC – Thr, ACA – Thr, ACG – ThrC
AAAU – Asn, AAC – Asn, AAA – Lys, AAG – LysA