A DIRICHLET PROCESS MIXTURE OF HIDDEN MARKOV MODELS FOR PROTEIN STRUCTURE PREDICTION.

A DIRICHLET PROCESS MIXTURE OF HIDDEN MARKOV MODELS FOR PROTEIN STRUCTURE PREDICTION.
复制标题

DOI:
10.1214/09-aoas296
复制
发表时间:
2010-06-01
期刊:
The annals of applied statistics
影响因子:
--
通讯作者:
Tsai JW
Tsai JW
中科院分区:
其他
文献类型:
--
作者:
Lennox KP;Dahl DB;Vannucci M;Day R;Tsai JW

文献摘要

被引文献

相似文献

通过对蛋白质扭转角的分布提供新的见解,最近的统计模型为这些数据指明了更有效的蛋白质结构预测方法。目前大多数方法都集中在一个单一的序列位置的双变量模型。然而,在蛋白质中的多个序列位置同时建模角度对具有相当大的价值。这种模型的一个应用领域是高度可变的环和转弯区域的结构预测。由于可用于估计这些扭转角分布的已知蛋白质结构的数量通常很小,因此这种建模是困难的。此外,数据是“稀疏的”,因为并非所有蛋白质在每个序列位置都具有角度对。我们提出了一个新的半参数模型的角度对在多个序列位置的联合分布。我们的模型通过利用蛋白质二级结构行为的已知信息来适应稀疏数据。我们证明了我们的技术,通过预测的扭转角在一个循环中的珠蛋白折叠家庭。我们的研究结果表明,基于模板的方法现在可以成功地扩展到建模出了名的困难的循环和转弯区域。
By providing new insights into the distribution of a protein’s torsion angles, recent statistical models for this data have pointed the way to more efficient methods for protein structure prediction. Most current approaches have concentrated on bivariate models at a single sequence position. There is, however, considerable value in simultaneously modeling angle pairs at multiple sequence positions in a protein. One area of application for such models is in structure prediction for the highly variable loop and turn regions. Such modeling is difficult due to the fact that the number of known protein structures available to estimate these torsion angle distributions is typically small. Furthermore, the data is “sparse” in that not all proteins have angle pairs at each sequence position. We propose a new semiparametric model for the joint distributions of angle pairs at multiple sequence positions. Our model accommodates sparse data by leveraging known information about the behavior of protein secondary structure. We demonstrate our technique by predicting the torsion angles in a loop from the globin fold family. Our results show that a template-based approach can now be successfully extended to modeling the notoriously difficult loop and turn regions.