Pieces of the puzzle: expressed sequence tags and the catalog of human genes

Pieces of the puzzle: expressed sequence tags and the catalog of human genes
复制标题

DOI:
10.1007/s001090050155
复制
发表时间:
1997-10-01
影响因子:
4.7
通讯作者:
Schuler, GD
Schuler, GD
中科院分区:
医学2区
文献类型:
--
作者:
Schuler, GD

文献摘要

被引文献

相似文献

想象一下,试图解决一个拼图游戏,而没有所有的碎片。这正是分子医学领域的研究人员在试图了解人类基因及其蛋白质产物如何相互作用以产生正常的生物功能时所面临的困境,这些功能如何在各种疾病状态下被破坏,以及如何通过分子干预恢复正常功能。这种对生命之谜的描述并不意味着否认环境和其他表观遗传因素的重要性,而只是为了定义一个谜题的边界,这个谜题的解决方案很容易在我们的掌握之中。为了进一步加深我们对人类生物学和遗传疾病遗传学的基本理解,编制一份完整的人类基因序列目录,并通过互联网向全世界的科学家提供这些信息,将是非常有价值的。在过去的几年里,与这个谜题相关的大量数据已经出现,但解决这个谜题仍然是生物信息学的挑战。在着手解决生命之谜之前,有一个粗略的概念,它包含多少块是有用的。换句话说,人类有多少基因?基于间接证据,估计约有64,000 [1]至80,000 [2]个基因。全基因组测序已被用于生成几种基因组相对较小的生物体的基因目录[3]。然而,人类基因组测序是一个更艰巨的任务,由于其巨大的规模(约30亿个碱基)。美国基因组计划始于1990年,其雄心勃勃的目标是在15年内(即到2005年)对人类基因组进行测序。不幸的是,只有大约2%的碱基组成了我们基因的蛋白质编码部分;其余98%的功能未知,通常被称为“垃圾DNA”。因此,基因组测序可能不是产生人类基因目录的最有效方法。许多研究人员主张对基因转录产物进行大规模测序,以互补DNA(cDNA)克隆的形式,作为整个人类基因组测序的前奏。正如布伦纳所说,“如果98%的基因组都是垃圾,那么最好的策略就是找到重要的2%,并首先对其进行测序。
Imagine trying to solve a jigsaw puzzle without having all of the pieces. This is exactly the dilemma faced by researchers in the field of molecular medicine when attempting to understand how human genes and their protein products interact with one another to lead to normal biological functions, how these functions can break down in various disease states, and how normal functions can be restored through molecular intervention. This description of the Puzzle of Life is not meant to deny the importance of environmental and other epigenetic factors, but is simply meant to define the boundaries of a puzzle whose solution is easily within our grasp. To further our basic understanding of human biology and the genetics of inherited diseases, it would be immensely valuable to compile a complete catalog of human gene sequences and to make this information available over the Internet to scientists around the world. Over the past few years huge amounts of data relevant to this puzzle have become available, but solving the puzzle remains a bioinformatics challenge.Before setting out to solve the Puzzle of Life, it would be useful to have a rough sense of how many pieces it contains. In other words, how many human genes are there? Based on indirect evidence, estimates ranging from approximately 64,000 [1] to 80,000 [2] genes have been advanced. Complete genomic sequencing has been used to generate gene catalogs for several organisms with relatively small genomes [3]. However, sequencing the human genome is a much more daunting task due to its immense size (about 3 billion bases). The United States Genome Project began in 1990 with the ambitious goal of sequencing the human genome within 15 years (ie, by the year 2005)[4]. Unfortunately, only about 2% of the total bases make up the protein-coding portions of our genes; the remaining 98% is of unknown function and often referred to as “junk DNA.” Thus, sequencing the genome may not be the most efficient way to generate a catalog of human genes. A number of investigators have advocated large-scale sequencing of the transcription products of genes, in the form of complimentary DNA (cDNA) clones, as a prelude to sequencing of the entire human genome. As Brenner [5] put it,“If something like 98% of the genome is junk, then the best strategy would be to find the important 2%, and sequence it first.”