Biology · Ch 5 — Molecular Basis of Inheritance
Human Genome Project and Rice Genome Project
Human Genome Project and Rice Genome Project
The Human Genome Project (HGP) was an ambitious, internationally coordinated scientific effort, formally launched in 1990 and largely completed by 2003, with the overarching goal of determining the complete DNA sequence of the entire human genome — all of the roughly three billion base pairs of DNA distributed across the 23 pairs of human chromosomes — and identifying and mapping the location of every gene contained within that sequence. The HGP was deliberately conceived and run as a "mega project," on a scale far larger than any biology project previously attempted, and its enormous cost was justified by an underlying principle its planners called cost-effective delivery: if the complete sequence and location of every human gene were known and made openly available in central databases, it would ultimately be far more cost-effective for future researchers, worldwide, to isolate, study and understand any particular gene of interest than for each research group to have to sequence and characterise it independently, from scratch.
The specific technical goals identified for the Human Genome Project included: identifying all of the estimated 20,000 to 25,000 genes present in human DNA; determining the complete sequence of the roughly 3 billion base pairs that make up the human genome; storing this vast quantity of sequence information in accessible, searchable central databases; developing faster, more efficient and lower-cost tools for DNA sequence analysis; transferring related sequencing and analysis technologies to the wider private sector and industry; and, importantly, also devoting explicit attention to the ethical, legal and social issues (ELSI) that such powerful, detailed genetic knowledge about human beings would inevitably raise.
Several of the HGP's key findings proved genuinely surprising to biologists at the time. The human genome contains a comparatively modest 20,000 to 25,000 protein-coding genes — a number that turned out to be far lower than many scientists had originally predicted (some estimates before the HGP had suggested well over 100,000 human genes), showing that human biological complexity does not simply scale in direct proportion to raw gene count. Protein-coding DNA sequences make up only a very small fraction — roughly about 2 percent — of the entire human genome; the remaining, much larger fraction of the genome consists of various forms of non-coding DNA, much of it once dismissively (and, we now understand, somewhat inaccurately) referred to as "junk DNA," though a great deal of this non-coding sequence is now known to serve important regulatory and structural functions. Repetitive sequences (stretches of DNA that repeat many times, sometimes millions of times, throughout the genome) make up a substantial proportion of the total human genome. The average human gene size is roughly 3,000 base pairs, though gene sizes vary enormously, with the largest known human gene (dystrophin) spanning some 2.4 million base pairs. Chromosome 1 is the largest human chromosome and carries the greatest number of genes (roughly 2,968 genes), while the Y chromosome is the smallest and carries by far the fewest. …