Learning generative models for protein fold families

Learning generative models for protein fold families
复制标题

DOI:
10.1002/prot.22934
复制
发表时间:
2011-04-01
影响因子:
2.9
通讯作者:
Langmead, Christopher James
Langmead, Christopher James
中科院分区:
生物学4区
文献类型:
--
作者:
Balakrishnan, Sivaraman;Kamisetty, Hetunandan;Langmead, Christopher James

文献摘要

被引文献

相似文献

我们介绍了一种新的方法来学习统计模型的多序列比对(MSA)的蛋白质。我们的方法,称为GREMLIN(蛋白质的生成规则化ModeLs),学习MSA内氨基酸组成的无向概率图形模型。所得到的模型编码了位置特异性保守统计和连续和长距离残基对之间的相关突变统计。用于从MSA学习图形模型的现有技术要么对MSA内的条件独立性做出强的且通常不适当的假设(例如,例如,在一个实施例中,隐马尔可夫模型),或者使用次优算法来学习模型的参数。与此相反,GREMLIN没有一个先验的假设条件的独立性内的MSA。我们制定并解决一个凸优化问题,从而保证我们找到一个全局最优的模型收敛。由此产生的模型也是生成的,允许设计与MSA中具有相同统计特性的新蛋白质序列。我们对广泛研究的WW和PDZ域的协变统计进行了详细的分析,并表明我们的方法优于现有的算法,用于从MSA学习无向概率图形模型。然后,我们将我们的方法应用于PFAM数据库中的另外71个家庭,并证明所得到的模型在预测准确性方面显着优于隐马尔可夫模型。
We introduce a new approach to learning statistical models from multiple sequence alignments (MSA) of proteins. Our method, called GREMLIN (Generative REgularized ModeLs of proteINs), learns an undirected probabilistic graphical model of the amino acid composition within the MSA. The resulting model encodes both the position- specific conservation statistics and the correlated mutation statistics between sequential and long-range pairs of residues. Existing techniques for learning graphical models from MSA either make strong, and often inappropriate assumptions about the conditional independencies within the MSA (e. g., Hidden Markov Models), or else use suboptimal algorithms to learn the parameters of the model. In contrast, GREMLIN makes no a priori assumptions about the conditional independencies within the MSA. We formulate and solve a convex optimization problem, thus guaranteeing that we find a globally optimal model at convergence. The resulting model is also generative, allowing for the design of new protein sequences that have the same statistical properties as those in the MSA. We perform a detailed analysis of covariation statistics on the extensively studied WW and PDZ domains and show that our method out-performs an existing algorithm for learning undirected probabilistic graphical models from MSA. We then apply our approach to 71 additional families from the PFAM database and demonstrate that the resulting models significantly out-perform Hidden Markov Models in terms of predictive accuracy.