Probabilistic methods of identifying genes in prokaryotic genomes: Connections to the FIMM theory

Probabilistic methods of identifying genes in prokaryotic genomes: Connections to the FIMM theory
复制标题

DOI:
10.1093/bib/5.2.118
复制
发表时间:
2004-06-01
影响因子:
9.5
通讯作者:
Borodovsky, M
Borodovsky, M
中科院分区:
生物学2区
文献类型:
--
作者:
Azad, RK;Borodovsky, M

文献摘要

被引文献

相似文献

在本文中,我们回顾了原核基因组中基因识别方法的发展,重点是与隐藏的马尔可夫模型(HMM)的一般理论的联系。我们表明,可以将基因标记(一种经常使用的基因调查工具)实现的贝叶斯方法增强并重新引入,作为HMM理论中描述的局部后解码的严格前向后退(FB)算法。另一个早期开发的方法,原核基因hmm,使用持续时间对HMM的viterbi算法进行修改,以识别给定DNA序列的隐藏功能状态的最可能的全局路径。 Genemark和Genemark.HMM程序值得一致地用于分析原核DNA序列,这些序列可以说不遵循任何确切的数学模型。使用FB算法的基因标记的新扩展已在软件程序genemark.fba中实现。鉴于DNA序列,该程序决定了每个核苷酸属于编码或非编码区域的后验概率。同样,对于任何开放阅读框架(ORF),它分配了一个分数定义为对所有路径的概率度量,这些路径穿过隐藏的状态,这些状态将ORF作为编码区域横穿ORF。在我们的测试中确定的genemark.FBA的预测准确性与初始(标准)基因计划的准确性进行了比较。与原核基因的比较。HMM还表明了原始基因检测的一定程度,但物种特异性的改进程度,即正确阅读框架(和终止密码子)。精确基因预测的准确性,它关注基因开始的精确预测(在原核生物基因组中明确定义了阅读框和终止密码子,因此,整个蛋白质产物)在基因标志中仍然更加准确,它使用了更精致的详细信息嗯,专门解决此任务。
In this paper, we review developments in probabilistic methods of gene recognition in prokaryotic genomes with the emphasis on connections to the general theory of hidden Markov models (HMM). We show that the Bayesian method implemented in GeneMark, a frequently used gene-finding tool, can be augmented and reintroduced as a rigorous forward-backward (FB) algorithm for local posterior decoding described in the HMM theory. Another earlier developed method, prokaryotic GeneMark.hmm, uses a modification of the Viterbi algorithm for HMM with duration to identify the most likely global path through hidden functional states given the DNA sequence. GeneMark and GeneMark.hmm programs are worth using in concert for analysing prokaryotic DNA sequences that arguably do not follow any exact mathematical model. The new extension of GeneMark using the FB algorithm was implemented in the software program GeneMark.fba. Given the DNA sequence, this program determines an a posteriori probability for each nucleotide to belong to coding or non-coding region. Also, for any open reading frame (ORF), it assigns a score defined as a probabilistic measure of all paths through hidden states that traverse the ORF as a coding region. The prediction accuracy of GeneMark.fba determined in our tests was compared favourably to the accuracy of the initial (standard) GeneMark program. Comparison to the prokaryotic GeneMark.hmm has also demonstrated a certain, yet species-specific, degree of improvement in raw gene detection, ie detection of correct reading frame (and stop codon). The accuracy of exact gene prediction, which is concerned about precise prediction of gene start (which in a prokaryotic genome unambiguously defines the reading frame and stop codon, thus, the whole protein product), still remains more accurate in GeneMarks, which uses more elaborate HMM to specifically address this task.