Microbial gene identification using interpolated Markov models

Microbial gene identification using interpolated Markov models
复制标题

DOI:
10.1093/nar/26.2.544
复制
发表时间:
1998-01-15
影响因子:
14.9
通讯作者:
White, O
White, O
中科院分区:
生物学2区
文献类型:
--
作者:
Salzberg, SL;Delcher, AL;White, O

文献摘要

被引文献

相似文献

本文介绍了一个新的系统,GLIMMER,用于寻找基因在微生物基因组中。在对流感嗜血杆菌、幽门螺杆菌和其他完整微生物基因组的一系列测试中,该系统已被证明在定位这些序列中的几乎所有基因方面非常准确,优于以前的方法。基于对幽门螺杆菌和流感螺杆菌的实验的保守估计是,该系统发现了>97%的所有基因,GLIMMER使用插值马尔可夫模型(IMM)作为捕获DNA序列中附近核苷酸之间依赖性的框架。基于IMM的方法基于可变上下文进行预测;即,一个可变长度的寡聚体在DNA序列中,由GLIMMER使用的上下文变化取决于序列的局部组成,因此,GLIMMER是更灵活和更强大的比固定顺序马尔可夫方法,这是以前的主要内容为基础的技术在微生物DNA中寻找基因。
This paper describes a new system, GLIMMER, for finding genes in microbial genomes. In a series of tests on Haemophilus influenzae, Helicobacter pylori and other complete microbial genomes, this system has proven to be very accurate at locating virtually ail the genes in these sequences, outperforming previous methods. A conservative estimate based on experiments on H.pylori and H.influenzae is that the system finds >97% of all genes, GLIMMER uses Interpolated Markov models (IMMs) as a framework for capturing dependencies between nearby nucleotides in a DNA sequence. An IMM-based method makes predictions based on a variable context; i.e., a variable-length oligomer in a DNA sequence, The context used by GLIMMER changes depending on the local composition of the sequence, As a result, GLIMMER is more flexible and more powerful than fixed-order Markov methods, which have previously been the primary content-based technique for finding genes in microbial DNA.