Efficient decoding algorithms for generalized hidden Markov model gene finders.

Efficient decoding algorithms for generalized hidden Markov model gene finders.
复制标题

DOI:
10.1186/1471-2105-6-16
复制
发表时间:
2005-01-24
期刊:
影响因子:
3
通讯作者:
Salzberg SL
Salzberg SL
中科院分区:
生物学4区
文献类型:
--
作者:
Majoros WH;Pertea M;Delcher AL;Salzberg SL

文献摘要

参考文献

被引文献

相似文献

广义隐马尔可夫模型(GHMM)已被证明是一个有用的框架,在真核生物基因组的计算基因预测的任务,由于其灵活性和概率基础。随着基因发现社区的焦点转向使用同源性信息来提高预测精度,正在探索基本GHMM模型的扩展作为将此同源性信息整合到预测过程中的可能方法。在这些扩展中特别突出的是那些要求同时预测两个或多个基因组中的基因的技术,从而显著增加了预测的计算成本,并突出了底层GHMM算法实现中速度和内存效率的重要性。不幸的是,实现一个有效的基于GHMM的基因发现者的任务已经是一个不平凡的任务,可以预料,随着我们的模型复杂性的增加,这项任务只会变得更加繁重。作为解决这些下一代系统的实施挑战的第一步,我们详细描述了两个基于GHMM的基因查找器的软件架构,一个包括常见的基于阵列的方法,和其他高度优化的算法,需要显着更少的内存,同时实现几乎相同的速度。然后,我们将展示如何通过优化内容传感器来加速这两种架构。最后,我们简要说明了这些优化对我们新的基于同源性的基因查找器TWAIN的可行性的影响。在描述了一些基于GHMM的基因搜索优化,并提供两个完整的开源软件系统体现这些方法,这是我们希望,其他人将能够更多地探索有前途的扩展GHMM框架,从而提高国家的最先进的基因预测技术。
The Generalized Hidden Markov Model (GHMM) has proven a useful framework for the task of computational gene prediction in eukaryotic genomes, due to its flexibility and probabilistic underpinnings. As the focus of the gene finding community shifts toward the use of homology information to improve prediction accuracy, extensions to the basic GHMM model are being explored as possible ways to integrate this homology information into the prediction process. Particularly prominent among these extensions are those techniques which call for the simultaneous prediction of genes in two or more genomes at once, thereby increasing significantly the computational cost of prediction and highlighting the importance of speed and memory efficiency in the implementation of the underlying GHMM algorithms. Unfortunately, the task of implementing an efficient GHMM-based gene finder is already a nontrivial one, and it can be expected that this task will only grow more onerous as our models increase in complexity. As a first step toward addressing the implementation challenges of these next-generation systems, we describe in detail two software architectures for GHMM-based gene finders, one comprising the common array-based approach, and the other a highly optimized algorithm which requires significantly less memory while achieving virtually identical speed. We then show how both of these architectures can be accelerated by a factor of two by optimizing their content sensors. We finish with a brief illustration of the impact these optimizations have had on the feasibility of our new homology-based gene finder, TWAIN. In describing a number of optimizations for GHMM-based gene finders and making available two complete open-source software systems embodying these methods, it is our hope that others will be more enabled to explore promising extensions to the GHMM framework, thereby improving the state-of-the-art in gene prediction techniques.
DOI: 10.1093/bioinformatics/btg1080
发表时间: 2003-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Stanke, Mario;Waack, Stephan
通讯作者: Waack, Stephan
DOI: 10.1002/0471250953.bia03as18
发表时间: 2007-06-01
影响因子: --
作者:
Schuster-Bockler, Benjamin;Bateman, Alex
通讯作者: Bateman, Alex
DOI: 10.1006/geno.1999.5854
发表时间: 1999-07-01
期刊: GENOMICS
影响因子: 4.4
作者:
Salzberg, SL;Pertea, M;Tettelin, H
通讯作者: Tettelin, H
DOI: 10.1101/gr.424203
发表时间: 2003-03-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Alexandersson, M;Cawley, S;Pachter, L
通讯作者: Pachter, L
DOI: 10.1006/jmbi.1997.0951
发表时间: 1997-04-25
影响因子: 5.6
作者:
Burge, C;Karlin, S
通讯作者: Karlin, S