Using the transcriptome to annotate the genome

Using the transcriptome to annotate the genome
复制标题

DOI:
10.1038/nbt0502-508
复制
发表时间:
2002-05-01
影响因子:
46.9
通讯作者:
Velculescu, VE
Velculescu, VE
中科院分区:
工程技术1区
文献类型:
--
作者:
Saha, S;Sparks, AB;Velculescu, VE

文献摘要

被引文献

相似文献

人类基因组计划的另一个挑战是对表达基因的识别和注释。公共和私人测序工作已经确定了15,000个符合基因严格标准的序列,例如与来自人类或其他物种的已知基因的对应关系,并做出了另一个类似的10,000- 20,000个置信度较低的基因预测,这些预测得到了各种类型的计算机证据的支持,包括同源性研究,域搜索和从头基因预测(1,2)。这些计算方法具有局限性,因为它们无法识别基因和外显子的显著部分,并且因为它们无法提供关于假设基因是否实际表达的明确证据(3,4)。由于计算机模拟方法识别的基因数量比预期的少(5-9),我们想知道高通量实验分析是否可以用于为假设基因的表达提供证据,并揭示以前未发现的基因。我们在这里描述了这种方法的发展-称为基因表达的长序列分析(LongSAGE),原始SAGE方法的改编(10)-可用于快速识别新基因和外显子。
A remaining challenge for the human genome project involves the identification and annotation of expressed genes. The public and private sequencing efforts have identified 15,000 sequences that meet stringent criteria for genes, such as correspondence with known genes from humans or other species, and have made another similar to10,000-20,000 gene predictions of lower confidence, supported by various types of in silico evidence, including homology studies, domain searches, and ab initio gene predictions(1,2). These computational methods have limitations, both because they are unable to identify a significant fraction of genes and exons and because they are unable to provide definitive evidence about whether a hypothetical gene is actually expressed(3,4). As the in silico approaches identified a smaller number of genes than anticipated(5-9), we wondered whether high-throughput experimental analyses could be used to provide evidence for the expression of hypothetical genes and to reveal previously undiscovered genes. We describe here the development of such a method-called long serial analysis of gene expression (LongSAGE), an adaption of the original SAGE approach(10)-that can be used to rapidly identify novel genes and exons.