Combining structural modeling with ensemble machine learning to accurately predict protein fold stability and binding affinity effects upon mutation.

Combining structural modeling with ensemble machine learning to accurately predict protein fold stability and binding affinity effects upon mutation.
复制标题

DOI:
10.1371/journal.pone.0107353
复制
发表时间:
2014
期刊:
影响因子:
3.7
通讯作者:
Kim PM
Kim PM
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Berliner N;Teyra J;Colak R;Garcia Lopez S;Kim PM

文献摘要

参考文献

被引文献

相似文献

测序方面的进步导致了突变的快速积累,其中一些与疾病有关。然而,为了得出机械论的结论,对这些突变的生物化学理解是必要的。对于编码突变,需要准确地预测蛋白质的稳定性或它们与结合伙伴的亲和力的显著变化。传统方法使用半经验力场,而新方法使用顺序和结构特征的机器学习。在这里,我们展示了如何将这两种方法结合起来,从而显著提高精度。我们介绍了ELASPIC,这是一种新的集成机器学习方法,能够预测结构域核心和结构域界面上突变对稳定性的影响。我们将半经验能量项、序列守恒和各种分子细节与随机梯度决策树(SGB-DT)算法相结合。我们的预测精度远远超过现有方法,稳定性相关系数达到0.77,亲和力预测相关系数达到0.75。值得注意的是,我们集成了同源建模来实现蛋白质组范围的预测,并表明对建模结构的准确预测是可能的。最后,ELASPIC显示了不同类型的疾病相关突变以及疾病和常见中性突变之间的显著差异。与试图预测突变的表型影响的纯基于序列的预测方法不同,我们的预测揭示了控制蛋白质不稳定性的分子细节,并帮助我们更好地理解疾病的分子原因。
Advances in sequencing have led to a rapid accumulation of mutations, some of which are associated with diseases. However, to draw mechanistic conclusions, a biochemical understanding of these mutations is necessary. For coding mutations, accurate prediction of significant changes in either the stability of proteins or their affinity to their binding partners is required. Traditional methods have used semi-empirical force fields, while newer methods employ machine learning of sequence and structural features. Here, we show how combining both of these approaches leads to a marked boost in accuracy. We introduce ELASPIC, a novel ensemble machine learning approach that is able to predict stability effects upon mutation in both, domain cores and domain-domain interfaces. We combine semi-empirical energy terms, sequence conservation, and a wide variety of molecular details with a Stochastic Gradient Boosting of Decision Trees (SGB-DT) algorithm. The accuracy of our predictions surpasses existing methods by a considerable margin, achieving correlation coefficients of 0.77 for stability, and 0.75 for affinity predictions. Notably, we integrated homology modeling to enable proteome-wide prediction and show that accurate prediction on modeled structures is possible. Lastly, ELASPIC showed significant differences between various types of disease-associated mutations, as well as between disease and common neutral mutations. Unlike pure sequence-based prediction methods that try to predict phenotypic effects of mutations, our predictions unravel the molecular details governing the protein instability, and help us better understand the molecular causes of diseases.
DOI: 10.1093/nar/gks1158
发表时间: 2013-01
影响因子: 14.9
作者:
Chatr-Aryamontri A;Breitkreutz BJ;Heinicke S;Boucher L;Winter A;Stark C;Nixon J;Ramage L;Kolas N;O'Donnell L;Reguly T;Breitkreutz A;Sellam A;Chen D;Chang C;Rust J;Livstone M;Oughtred R;Dolinski K;Tyers M
通讯作者: Tyers M
DOI: 10.1093/nar/gkq929
发表时间: 2011-01
影响因子: 14.9
作者:
Forbes SA;Bindal N;Bamford S;Cole C;Kok CY;Beare D;Jia M;Shepherd R;Leung K;Menzies A;Teague JW;Campbell PJ;Stratton MR;Futreal PA
通讯作者: Futreal PA
DOI: 10.1093/nar/gkr996
发表时间: 2012-01
影响因子: 14.9
作者:
De Baets G;Van Durme J;Reumers J;Maurer-Stroh S;Vanhee P;Dopazo J;Schymkowitz J;Rousseau F
通讯作者: Rousseau F
DOI: 10.1093/nar/gkp985
发表时间: 2010-01
影响因子: 14.9
作者:
Finn RD;Mistry J;Tate J;Coggill P;Heger A;Pollington JE;Gavin OL;Gunasekaran P;Ceric G;Forslund K;Holm L;Sonnhammer EL;Eddy SR;Bateman A
通讯作者: Bateman A
DOI: 10.1093/bioinformatics/bti1109
发表时间: 2005-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Capriotti, E;Fariselli, P;Casadio, R
通讯作者: Casadio, R