Skip to content

Biology · Ch 5 — Molecular Basis of Inheritance

The Genetic Code

5.13

The Genetic Code

Once a gene's information has been transcribed into mRNA, the cell faces a further, purely informational problem: mRNA is written in an alphabet of just four different nucleotide bases, while a protein is built from twenty different amino acids — so some fixed, systematic rule (a "code") must specify exactly how a sequence of nucleotide bases in mRNA is translated into a corresponding sequence of amino acids in a protein. This rule is called the genetic code. The genetic code was worked out experimentally through the pioneering biochemical work of Har Gobind Khorana and Marshall Nirenberg (along with several other contributing scientists) during the 1960s, using synthetic RNA molecules of known, defined sequence and testing which amino acid each such sequence directed a cell-free translation system to incorporate.

Because there are only four different bases available but twenty different amino acids that must each be specified, a code based on single bases (which could specify only 4 possibilities) or even on pairs of bases (4×4 = 16 possibilities) would not provide enough distinct combinations to specify all twenty amino acids uniquely. A code based on a group of three consecutive bases, however, gives 4×4×4 = 64 possible combinations — comfortably more than the twenty amino acids that must be encoded — and this is exactly the system living cells use: each group of three consecutive nucleotide bases in mRNA, called a codon, specifies one particular amino acid (or, for a small number of codons, a stop signal rather than an amino acid).

The genetic code has several well-characterised general properties. It is a triplet code, meaning each codon consists of exactly three consecutive bases. It is degenerate (or redundant): since 64 possible codons must specify only 20 amino acids, most amino acids are actually specified by more than one different codon (for example, several different codons all specify the amino acid leucine) — this degeneracy provides the cell with some built-in tolerance against certain mutations, since a base change that produces a different but "synonymous" codon may still specify the identical amino acid. The code is unambiguous, meaning that any one specific codon always specifies exactly the same single amino acid, never two or more different amino acids (degeneracy runs only in one direction: many codons can specify one amino acid, but one codon never specifies more than one amino acid). The code is nearly universal, meaning the same codon specifies the same amino acid across almost all known living organisms, from bacteria to plants to animals — a powerful piece of evidence for the common evolutionary origin of all life on Earth (with only a small number of documented exceptions, mainly in mitochondrial genetic codes). The code is also commaless and non-overlapping, meaning that codons are read consecutively from a fixed starting point, in one direction, three bases at a time, with no "punctuation" bases skipped between successive codons and no base shared between two different, overlapping codons. …

Table 5.2Key Properties of the Genetic Code
PropertyWhat it meansExample / evidence
TripletEach codon = 3 consecutive bases4x4x4 = 64 codons for 20 amino acids
Degenerate (redundant)Most amino acids specified by >1 codonLeucine has several synonymous codons
UnambiguousOne codon never specifies more than one amino acidDegeneracy runs only one way
(Near-)UniversalSame codon = same amino acid in almost all organismsEvidence for common evolutionary origin