Combining phylogenetic and hidden Markov models in biosequence analysis

Combining phylogenetic and hidden Markov models in biosequence analysis
复制标题

DOI:
10.1089/1066527041410472
复制
发表时间:
2004-01-01
影响因子:
1.7
通讯作者:
Haussler, D
Haussler, D
中科院分区:
生物学4区
文献类型:
--
作者:
Siepel, A;Haussler, D

文献摘要

被引文献

相似文献

近年来已经出现了一些模型,这些模型不仅考虑了基因组每个地点的进化历史的替代方式,而且还考虑了从一个站点变为另一个站点的过程的方式。这些模型结合了分子进化的系统发育模型,这些模型适用于各个位点,以及隐藏的马尔可夫模型,这些模型允许从站点到站点变化。除了改善普通系统发育模型的现实主义外,它们可能是推理和预测示例的非常强大的工具,用于基因查找或预测二级结构。在本文中,我们回顾了合并的系统发育和隐藏马尔可夫模型的进度,并为先前的工作提供了一些扩展。我们的主要结果是一种简单有效的方法,用于适应HMM中高阶状态的方法,该方法允许替代的上下文依赖性模型,即考虑到相邻基础对替代模式的影响的模型。我们提出了实验结果,表明高阶状态,自相关速率和多个功能类别都可以显着改善系统发育和隐藏的马尔可夫模型的拟合度,而高阶状态的影响特别明显。
A few models have appeared in recent years that consider not only the way substitutions occur through evolutionary history at each site of a genome, but also the way the process changes from one site to the next. These models combine phylogenetic models of molecular evolution, which apply to individual sites, and hidden Markov models, which allow for changes from site to site. Besides improving the realism of ordinary phylogenetic models, they are potentially very powerful tools for inference and prediction-for example, for gene finding or prediction of secondary structure. In this paper, we review progress on combined phylogenetic and hidden Markov models and present some extensions to previous work. Our main result is a simple and efficient method for accommodating higher-order states in the HMM, which allows for context-dependent models of substitution-that is, models that consider the effects of neighboring bases on the pattern of substitution. We present experimental results indicating that higher-order states, autocorrelated rates, and multiple functional categories all lead to significant improvements in the fit of a combined phylogenetic and hidden Markov model, with the effect of higher-order states being particularly pronounced.