Skip to content

Botany · Ch 7 — Molecular Basis of Inheritance

Human Genome Project

7.9

Human Genome Project

The Human Genome Project (HGP) was a landmark international scientific research project with one primary goal: to determine the complete sequence of DNA base pairs that make up the human genome, and to identify and map all the genes it contains. It is the largest biological project ever undertaken.

The project was formally begun in 1990 and was declared complete in 2003, two years ahead of schedule. It was coordinated by the U.S. Department of Energy and the National Institutes of Health, with major contributions from laboratories in the United Kingdom, France, Germany, Japan, China, and other countries.

Goals of the HGP

The project had several clearly defined goals:

  • Identify all the approximately 20,000-25,000 genes in human DNA.
  • Determine the sequences of the 3 billion chemical base pairs that make up human DNA.
  • Store this information in databases for public access.
  • Improve tools for data analysis.
  • Transfer related technologies to the private sector.
  • Address the ethical, legal, and social issues (ELSI) that may arise from the project.

Methodology

Two main approaches were used to sequence the genome:

  1. Expressed Sequence Tags (ESTs): This method focused on sequencing only the parts of the genome that are actually expressed as proteins (the coding regions, or genes). It was a faster way to identify genes.
  2. Sequence Annotation: This was the more comprehensive approach. The entire genome was first sequenced, and then the sequence was analysed to identify the locations of genes, regulatory sequences, and other functional elements.

The actual sequencing work was done using automated DNA sequencers. The DNA was first broken into small fragments, each fragment was sequenced, and then powerful computers were used to assemble these overlapping fragments back into the correct order — a process often compared to putting together a giant jigsaw puzzle.

Key Findings of the Human Genome Project

The HGP produced several surprising and important results:

  • Gene Count: The human genome contains only about 30,000 genes. This was far fewer than scientists had expected (earlier estimates ranged from 80,000 to 1,40,000). This number is only about twice that of a fruit fly or a roundworm.
  • Repetitive DNA: Over 50% of the human genome is made up of non-coding, repetitive DNA sequences. These sequences do not code for proteins, and their function was (and in many cases still is) not fully understood.
  • Largest Gene: The largest known human gene is that for dystrophin, a muscle protein. It is about 2.4 million base pairs long.
  • Chromosome 1: Chromosome 1 has the most genes (approximately 2,968), while the Y chromosome has the fewest (approximately 231).
  • Protein Similarity: The sequences of human proteins are remarkably similar to those of other organisms. For example, about 99.9% of the DNA sequence is identical between any two humans. The vast majority of human genes have counterparts in other animals, including mice and even fruit flies. …
Figure 5.15A representative diagram of human genome project
Fig. 5.15 — A representative diagram of human genome project

Drawn by us to help you understand the concept clearly, and verified to make sure it's accurate. For exams, practice from your textbook's own diagram.

The figure is a flowchart-style diagram that walks through the key steps of the Human Genome Project’s sequencing strategy. It begins at the top with a double-stranded DNA molecule — this represents the entire human genome, which is far too large to sequence in one piece.

The first step shown is fragmentation: the DNA is cut into smaller, overlapping fragments. These fragments are then cloned into vectors — the diagram specifically labels two types: BAC (Bacterial Artificial Chromosome) and YAC (Yeast Artificial Chromosome). Each fragment is inserted into one of these vectors, creating a library of cloned DNA pieces. The figure likely shows these as small circular or linear representations of the vector carrying a human DNA insert.

Next, each cloned fragment is sequenced individually. The diagram probably depicts short stretches of nucleotide sequence (A, T, G, C) emerging from each clone. This is the raw data stage.

The critical step is assembly by computer. The figure shows how the computer uses the overlapping regions at the ends of the sequenced fragments to align them. Arrows or connecting lines likely indicate how the overlapping sequences match up, allowing the computer to stitch the fragments together in the correct order. The final result at the bottom of the diagram is a continuous, assembled DNA sequence — the complete genome. …