Analysis of an optimal hidden Markov model for secondary structure prediction.

Analysis of an optimal hidden Markov model for secondary structure prediction.
复制标题

分析用于二级结构预测的最佳隐藏马尔可夫模型。

DOI:
10.1186/1472-6807-6-25
复制
发表时间:
2006-12-13
影响因子:
--
通讯作者:
Rodolphe, Francois
Rodolphe, Francois
中科院分区:
生物4区
文献类型:
--
作者:
Martin, Juliette;Gibrat, Jean-Francois;Rodolphe, Francois

文献摘要

被引文献

相似文献

二级结构预测是 3D 结构预测的有用的第一步。许多成功的二级结构预测方法都使用神经网络,但不幸的是,神经网络不能直观地解释。相反,隐马尔可夫模型是图形可解释模型。此外,它们已成功用于许多生物信息学应用。因为它们提供了强大的统计背景并允许模型解释,所以我们提出了一种基于隐马尔可夫模型的方法。我们的 HMM 是在没有先验知识的情况下设计的。它是使用统计和准确性标准在尺寸不断增加的模型集合中选择的。生成的模型有 36 个隐藏状态:15 个对 α 螺旋建模,12 个对线圈建模,9 个对 β 链建模。隐藏状态和状态发射概率之间的联系反映了蛋白质结构到二级结构片段的组织。我们首先分析模型特征,看看它如何提供局部结构的新视野。然后我们将其用于二级结构预测。我们的模型似乎对单个序列非常有效,Q3 得分为 68.8%,比 PSIPRED 对单个序列的预测高出一分多。该方法的直接扩展允许使用多个序列比对,将 Q3 分数提高到 75.5%。这里提出的隐马尔可夫模型仅使用有限数量的参数即可实现有价值的预测结果。它为蛋白质二级结构提供了一个可解释的框架。此外,它可以用作生成具有给定二级结构内容的蛋白质序列的工具。
Secondary structure prediction is a useful first step toward 3D structure prediction. A number of successful secondary structure prediction methods use neural networks, but unfortunately, neural networks are not intuitively interpretable. On the contrary, hidden Markov models are graphical interpretable models. Moreover, they have been successfully used in many bioinformatic applications. Because they offer a strong statistical background and allow model interpretation, we propose a method based on hidden Markov models. Our HMM is designed without prior knowledge. It is chosen within a collection of models of increasing size, using statistical and accuracy criteria. The resulting model has 36 hidden states: 15 that model α-helices, 12 that model coil and 9 that model β-strands. Connections between hidden states and state emission probabilities reflect the organization of protein structures into secondary structure segments. We start by analyzing the model features and see how it offers a new vision of local structures. We then use it for secondary structure prediction. Our model appears to be very efficient on single sequences, with a Q3 score of 68.8%, more than one point above PSIPRED prediction on single sequences. A straightforward extension of the method allows the use of multiple sequence alignments, rising the Q3 score to 75.5%. The hidden Markov model presented here achieves valuable prediction results using only a limited number of parameters. It provides an interpretable framework for protein secondary structure architecture. Furthermore, it can be used as a tool for generating protein sequences with a given secondary structure content.