Computational methods for the identification of genes in vertebrate genomic sequences

Computational methods for the identification of genes in vertebrate genomic sequences
复制标题

DOI:
10.1093/hmg/6.10.1735
复制
发表时间:
1997-01-01
影响因子:
3.5
通讯作者:
Claverie, JM
Claverie, JM
中科院分区:
生物学2区
文献类型:
--
作者:
Claverie, JM

文献摘要

被引文献

相似文献

在匿名基因组序列中识别基因的新方法的研究已经持续了15年以上,在这段时间里,该领域已经从设计程序来识别紧凑线粒体或细菌基因组中的蛋白质编码区,到预测脊椎动物多外显子基因的详细组织结构。目前可用的最好的程序完美地定位了超过80%的内部编码外显子,并且只有5%的预测不与真实的外显子重叠,考虑到这样的准确性,计算方法确实非常有用;然而,它们并没有减少对实验验证的需要。如果对基因的编码部分(内部编码外显子)的识别效果令人满意,转录本的全长(基因的5‘和3’端)的确定和启动子区域的定位仍然是不可靠的,随着人类和小鼠基因组测序计划进入生产模式,对百万碱基长的匿名基因组序列的全自动注释是生物信息学的下一个重大挑战。
Research into new methods to identify genes in anonymous genomic sequences has been going on for more than 15 years, Over this period of time, the field has evolved from the designing of programs to identify protein coding regions in compact mitochondrial or bacterial genomes, to the challenge of predicting the detailed organization of multi-exon vertebrate genes. The best program currently available perfectly locates more than 80% of the internal coding exons, and only 5% of the predictions do not overlap a real exon, Given such accuracy, computational methods are indeed very useful; however, they do not alleviate the need for experimental validation. If the performances are satisfactory for the identification of the coding moiety of genes (internal coding exons), the determination of the full extent of the transcript (5' and 3' extremities of the gene) and the location of promoter regions are still unreliable, As the human and mouse genome sequencing projects enter a production mode, the fully automated annotation of megabase-long anonymous genomic sequences is the next big challenge in bioinformatics.