Prediction of hot spot residues at protein-protein interfaces by combining machine learning and energy-based methods.

Prediction of hot spot residues at protein-protein interfaces by combining machine learning and energy-based methods.
复制标题

DOI:
10.1186/1471-2105-10-365
复制
发表时间:
2009-10-30
期刊:
影响因子:
3
通讯作者:
Jones DT
Jones DT
中科院分区:
生物学4区
文献类型:
--
作者:
Lise S;Archambeau C;Pontil M;Jones DT

文献摘要

参考文献

被引文献

相似文献

丙氨酸扫描诱变是研究蛋白质复合体结构和能量特征的一种强有力的实验方法。单个氨基酸被系统地突变为丙氨酸,并测量结合自由能(ΔΔG)的变化。几个实验表明,蛋白质-蛋白质的相互作用严重依赖于界面上的几个残基(“热点”)。热点对结合自由能的贡献占主导地位,如果它们发生突变,可能会破坏相互作用。由于突变研究需要大量的实验工作,因此需要准确可靠的计算方法。这样的方法还将增加我们对蛋白质-蛋白质识别中亲和力和特异性决定因素的理解。我们提出了一种新的计算策略来识别热点残基,给出了一个复合体的结构。我们考虑了影响热点相互作用的基本能量项,即范德华势、溶剂化能、氢键和库仑静电学。我们将它们作为输入特征,并基于一组丙氨酸突变的训练实例,使用支持向量机和高斯过程等机器学习算法对它们进行优化组合和集成。实验结果表明,该方法在预测热点方面是有效的,并且优于其他现有的预测方法。特别是,我们发现使用半监督学习方案--转导支持向量机的性能最好。当热点被定义为ΔΔG≥为2千卡/摩尔的残基时,我们的方法获得了56%的准确率和65%的召回率。我们开发了一种混合方案,其中能量项被用作机器学习模型的输入特征。该策略结合了机器学习和基于能量的方法的优点。虽然到目前为止,这两种方法主要是单独应用于生物分子问题,但我们的调查结果表明,它们的结合将获得实质性的好处。
Alanine scanning mutagenesis is a powerful experimental methodology for investigating the structural and energetic characteristics of protein complexes. Individual amino-acids are systematically mutated to alanine and changes in free energy of binding (ΔΔG) measured. Several experiments have shown that protein-protein interactions are critically dependent on just a few residues ("hot spots") at the interface. Hot spots make a dominant contribution to the free energy of binding and if mutated they can disrupt the interaction. As mutagenesis studies require significant experimental efforts, there is a need for accurate and reliable computational methods. Such methods would also add to our understanding of the determinants of affinity and specificity in protein-protein recognition. We present a novel computational strategy to identify hot spot residues, given the structure of a complex. We consider the basic energetic terms that contribute to hot spot interactions, i.e. van der Waals potentials, solvation energy, hydrogen bonds and Coulomb electrostatics. We treat them as input features and use machine learning algorithms such as Support Vector Machines and Gaussian Processes to optimally combine and integrate them, based on a set of training examples of alanine mutations. We show that our approach is effective in predicting hot spots and it compares favourably to other available methods. In particular we find the best performances using Transductive Support Vector Machines, a semi-supervised learning scheme. When hot spots are defined as those residues for which ΔΔG ≥ 2 kcal/mol, our method achieves a precision and a recall respectively of 56% and 65%. We have developed an hybrid scheme in which energy terms are used as input features of machine learning models. This strategy combines the strengths of machine learning and energy-based methods. Although so far these two types of approaches have mainly been applied separately to biomolecular problems, the results of our investigation indicate that there are substantial benefits to be gained by their integration.
DOI: 10.1016/s0022-2836(02)00442-4
发表时间: 2002-07-05
影响因子: 5.6
作者:
Guerois, R;Nielsen, JE;Serrano, L
通讯作者: Serrano, L
DOI: 10.1002/jcc.20893
发表时间: 2008-06-01
影响因子: 3
作者:
Kim, Ryangguk;Skolnick, Jeffrey
通讯作者: Skolnick, Jeffrey
DOI: 10.1186/1471-2105-9-447
发表时间: 2008-10-21
期刊: BMC bioinformatics
影响因子: 3
作者:
Grosdidier S;Fernández-Recio J
通讯作者: Fernández-Recio J
DOI: 10.1006/jmbi.2001.5009
发表时间: 2001-09-28
影响因子: 5.6
作者:
Elcock, AH
通讯作者: Elcock, AH
DOI: 10.1093/bioinformatics/bti1109
发表时间: 2005-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Capriotti, E;Fariselli, P;Casadio, R
通讯作者: Casadio, R