Genome annotation past, present, and future: How to define an ORF at each locus

Genome annotation past, present, and future: How to define an ORF at each locus
复制标题

DOI:
10.1101/gr.3866105
复制
发表时间:
2005-12-01
期刊:
影响因子:
7
通讯作者:
Brent, MR
Brent, MR
中科院分区:
生物学1区
文献类型:
--
作者:
Brent, MR

文献摘要

被引文献

相似文献

在竞争、自动化和技术的推动下,基因组学界已经远远超过了在2005年前完成人类基因组测序的目标。通过分析哺乳动物基因组,我们揭示了我们的DNA序列的历史,确定了选择性剪接的RNA和逆转录的假基因非常丰富,并瞥见了在基因调控中发挥重要作用的大量非编码RNA。最终,基因组科学可能会提供这些元素的全面目录。然而,我们在过去10年的大部分时间里一直使用的方法甚至不能为每个基因产生一个完整的开放阅读框架(CRF)--这是迈向全面目录的漫长攀登中的第一个平台。这些策略-随机选择的cDNA克隆测序,在其他生物体中鉴定的蛋白质序列比对,测序更多的基因组,和人工治疗-将不得不通过大规模扩增和特定的预测mRNA测序来补充。在过去的10年里,基因预测的稳步改进提高了这种方法的效率,降低了成本。在这个观点中,我回顾了大约10年前的基因预测状况,总结了自那时以来所取得的进展,认为我们迄今为止所依赖的主要ORF鉴定方法是不够的,并建议完成蛋白质编码基因目录1.0版的路径。
Driven by competition, automation, and technology, the genomics community has far exceeded its ambition to sequence the human genome by 2005. By analyzing mammalian genomes, we have shed light on the history of our DNA sequence, determined that alternatively spliced RNAs and retroposed pseudogenes are incredibly abundant, and glimpsed the apparently huge number of non-coding RNAs that play significant roles in gene regulation. Ultimately, genome science is likely to provide comprehensive catalogs of these elements. However, the methods we have been using for most of the last 10 years will not yield even one complete open reading frame (CRF) for every gene-the first plateau on the long climb toward a comprehensive catalog. These strategies-sequencing randomly selected cDNA clones, aligning protein sequences identified in other organisms, sequencing more genomes, and manual curation-will have to be supplemented by large-scale amplification and sequencing of specific predicted mRNAs. The steady improvements in gene prediction that have occurred over the last 10 years have increased the efficacy of this approach and decreased its cost. In this Perspective, I review the state of gene prediction roughly 10 years ago, summarize the progress that has been made since, argue that the primary ORF identification methods we have relied on so far are inadequate, and recommend a path toward completing the Catalog of Protein Coding Genes, Version 1.0.