GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions

GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions
复制标题

DOI:
10.1093/nar/29.12.2607
复制
发表时间:
2001-06-15
影响因子:
14.9
通讯作者:
Borodovsky, M
Borodovsky, M
中科院分区:
生物学2区
文献类型:
--
作者:
Besemer, J;Lomsadze, A;Borodovsky, M

文献摘要

被引文献

相似文献

提高基因起始点预测的准确性是计算机预测原核基因的几个未决问题之一。它的困难是由于缺乏相对较强的序列模式识别真正的翻译起始位点。在目前的文件中,我们表明,基因启动预测的准确性可以提高通过结合模型的蛋白质编码和非编码区和模型的调控位点附近的基因启动内的迭代隐马尔可夫模型为基础的算法。新的基因预测方法称为GeneMarkS,利用非监督训练程序,可用于新测序的原核基因组,而无需任何蛋白质或rRNA基因的先验知识。GeneMarkS的实现使用了基因发现程序GeneMark.hmm的改进版本、编码和非编码区域的启发式马尔可夫模型以及Gibbs抽样多重比对程序。GeneMarkS精确地预测了GenBank注释的枯草芽孢杆菌基因的83.2%的翻译起始和实验验证的大肠杆菌基因组中的94.4%的翻译起始。我们还观察到,GeneMarkS在鉴定含有真实的基因的开放阅读框方面检测原核基因,其准确度与目前使用的最佳基因检测方法的水平相匹配,准确的翻译起始预测,除了蛋白质序列N-末端数据的细化之外,还提供了位于基因起始上游的序列区域的精确定位的益处。因此,可以更高精度地揭示和分析与转录和翻译调控位点相关的序列基序。这些图案被证明具有显着的变异性,其中的功能和进化的连接进行了讨论。
Improving the accuracy of prediction of gene starts is one of a few remaining open problems in computer prediction of prokaryotic genes. Its difficulty is caused by the absence of relatively strong sequence patterns identifying true translation initiation sites. In the current paper we show that the accuracy of gene start prediction can be improved by combining models of protein-coding and non-coding regions and models of regulatory sites near gene start within an iterative Hidden Markov model based algorithm. The new gene prediction method, called GeneMarkS, utilizes a non-supervised training procedure and can be used for a newly sequenced prokaryotic genome with no prior knowledge of any protein or rRNA genes. The GeneMarkS implementation uses an improved version of the gene finding program GeneMark.hmm, heuristic Markov models of coding and non-coding regions and the Gibbs sampling multiple alignment program. GeneMarkS predicted precisely 83.2% of the translation starts of GenBank annotated Bacillus subtilis genes and 94.4% of translation starts in an experimentally validated set of Escherichia coli genes, We have also observed that GeneMarkS detects prokaryotic genes, in terms of identifying open reading frames containing real genes, with an accuracy matching the level of the best currently used gene detection methods, Accurate translation start prediction, in addition to the refinement of protein sequence N-terminal data, provides the benefit of precise positioning of the sequence region situated upstream to a gene start. Therefore, sequence motifs related to transcription and translation regulatory sites can be revealed and analyzed with higher precision. These motifs were shown to possess a significant variability, the functional and evolutionary connections of which are discussed.