Heuristic approach to deriving models for gene finding

Heuristic approach to deriving models for gene finding
复制标题

DOI:
10.1093/nar/27.19.3911
复制
发表时间:
1999-10-01
影响因子:
14.9
通讯作者:
Borodovsky, M
Borodovsky, M
中科院分区:
生物学2区
文献类型:
--
作者:
Besemer, J;Borodovsky, M

文献摘要

被引文献

相似文献

在DNA序列中精确寻找基因的计算机方法需要蛋白质编码区和非编码区的模型,这些模型要么来自实验验证的训练集,要么来自大量的匿名DNA序列。在这里,我们提出了一种新的启发式方法,生成相当准确的蛋白质编码区非齐次马尔可夫模型。这种新方法只需要少量的DNA序列数据,就可以通过网络服务器为任何>400 nt的DNA序列“即时”建立模型。使用GeneMark.hmm程序对10个完整的细菌基因组进行的测试表明,新模型平均能够检测93.1%的注释基因,而传统训练构建的模型平均能够预测93.9%的基因。通过启发式方法建立的模型可以用于在匿名原核基因组的小片段中以及在细胞器、病毒、质粒和质粒的基因组中找到基因,以及在高度不均匀的基因组中需要根据局部DNA组成调整模型。启发式方法也给出了深入了解密码子使用模式的演变机制。
Computer methods of accurate gene finding in DNA sequences require models of protein coding and non-coding regions derived either from experimentally validated training sets or from large amounts of anonymous DNA sequence. Here we propose a new, heuristic method producing fairly accurate inhomogeneous Markov models of protein coding regions. The new method needs such a small amount of DNA sequence data that the model can be built 'on the fly' by a web server for any DNA sequence >400 nt. Tests on 10 complete bacterial genomes performed with the GeneMark.hmm program demonstrated the ability of the new models to detect 93.1% of annotated genes on average, while models built by traditional training predict an average of 93.9% of genes. Models built by the heuristic approach could be used to find genes in small fragments of anonymous prokaryotic genomes and in genomes of organelles, viruses, phages and plasmids, as well as in highly inhomogeneous genomes where adjustment of models to local DNA composition is needed. The heuristic method also gives an insight into the mechanism of codon usage pattern evolution.