Markovian structures in biological sequence alignments

Markovian structures in biological sequence alignments
复制标题

DOI:
10.2307/2669673
复制
发表时间:
1999-03-01
影响因子:
3.7
通讯作者:
Lawrence, CE
Lawrence, CE
中科院分区:
数学1区
文献类型:
--
作者:
Liu, JS;Neuwald, AF;Lawrence, CE

文献摘要

被引文献

相似文献

多个同源生物聚合物序列的比对在蛋白质建模和工程、分子进化以及基因功能和基因产物结构的预测方面的研究中是至关重要的。在这篇文章中,我们提供了一个连贯的观点,最近的两个模型用于多序列的隐马尔可夫模型(HMM)和基于块的主题模型,开发一套新的算法,既有灵敏度的基于块的模型和HMM的灵活性。特别是,我们将标准HMM分解为两个组件:插入组件,这是由所谓的“传播模型”捕获,和删除组件,这是由删除向量描述。这种分解作为生物特异性和模型灵活性之间的合理妥协的基础。此外,我们引入了贝叶斯模型的选择标准,结合传播模型,遗传算法,和其他计算方面,形成了核心的PROBE,多对齐和数据库搜索方法。我们的方法的应用程序的GTdR家族的蛋白质序列产生的比对,通过与已知的三级结构比较确认。
The alignment of multiple homologous biopolymer sequences is crucial in research on protein modeling and engineering, molecular evolution, and prediction in terms of both gene function and gene product structure. In this article we provide a coherent view of the two recent models used for multiple sequence alignment-the hidden Markov model (HMM) and the block-based motif model-to develop a set of new algorithms that have both the sensitivity of the block-based model and the flexibility of the HMM. In particular, we decompose the standard HMM into two components: the insertion component, which is captured by the so-called "propagation model," and the deletion component, which is described by a deletion vector. Such a decomposition serves as a basis for rational compromise between biological specificity and model flexibility. Furthermore, we introduce a Bayesian model selection criterion that-in combination with the propagation model, genetic algorithm, and other computational aspects-forms the core of PROBE, a multiple alignment and database search methodology. The application of our method to a GTPase family of protein sequences yields an alignment that is confirmed by comparison with known tertiary structures.