Zoology · Ch 5 — Molecular Genetics
Goals and Methodologies of Human Genome Project
Goals and Methodologies of Human Genome Project
The Human Genome Project pursued six explicit goals: to identify every one of the roughly 30,000 genes present in human DNA; to determine the precise sequence of the three billion chemical base pairs that make up the entire human genome; to store all of this sequence information systematically in accessible databases; to improve the computational tools available for analysing genomic data; to transfer the resulting technologies and know-how to other sectors, including industry; and to proactively address the ethical, legal and social issues, abbreviated ELSI, that this new genetic knowledge could reasonably be expected to raise. To reach these goals, the project relied on two complementary methodological approaches. The first approach concentrated specifically on identifying every gene that is actually expressed as RNA, using a technique called Expressed Sequence Tags (ESTs), which catalogues short fragments of expressed genes directly. The second, broader approach was whole-genome sequence annotation: rather than focusing only on expressed genes, this approach sequenced the entire genome, including every coding and non-coding stretch of DNA, and only afterward went back and assigned likely biological functions to different regions of the completed sequence. For the sequencing work itself, total genomic DNA extracted from a cell was cut into random, more manageable fragments and cloned into specialised host organisms, most often bacteria or yeast, carried on purpose-built cloning vectors known as Bacterial Artificial Chromosomes (BACs) and Yeast Artificial Chromosomes (YACs) respectively; this cloning step amplified each fragment into many identical copies, making it practical to sequence. The amplified fragments were then read out using automated DNA sequencing machines, a technology originally developed by Frederick Sanger, and the resulting short sequence reads were computationally reassembled into a single, continuous genome sequence by identifying and aligning the short regions where neighbouring fragments happened to overlap. Once assembled, the sequences were annotated, systematically labelled with likely gene identities and functions, and assigned to their correct chromosome, a process aided by mapping known variation in restriction-endonuclea …