Prediction of interface residues in protein-protein complexes by a consensus neural network method: Test against NMR data

Prediction of interface residues in protein-protein complexes by a consensus neural network method: Test against NMR data
复制标题

DOI:
10.1002/prot.20514
复制
发表时间:
2005-10-01
影响因子:
2.9
通讯作者:
Zhou, HX
Zhou, HX
中科院分区:
生物学4区
文献类型:
--
作者:
Chen, HL;Zhou, HX

文献摘要

被引文献

相似文献

存放在蛋白质数据库中的蛋白质-蛋白质复合物结构的数量正在迅速增长。这些结构为预测新蛋白质复合物的结构提供了重要信息。这促使我们开发PPISP方法来预测蛋白质-蛋白质复合物中的界面残基。在PPISP中,空间相邻表面残基的序列分布和溶剂可及性被用作神经网络的输入。该网络在从蛋白质数据库收集的天然界面残基上进行训练。当时的预测准确率为70%,天然界面残基覆盖率为47%。现在我们已经广泛地改进了PPISP。训练集现在由1156条非同源蛋白质链组成。对100条非同源蛋白质链的测试表明,预测准确率提高到80%,覆盖率为51%。为了解决与单个神经网络模型相关的过度预测和预测不足的问题,我们开发了一种共识方法,该方法将来自多个模型的预测与不同级别的准确性和覆盖率相结合。应用于68种蛋白质的蛋白质对接基准集,共识方法在准确性上优于最佳个体模型3-8个百分点。为了证明cons-PPISP的预测能力,测试了具有通过NMR表征的界面的8种复合物形成蛋白。这些蛋白质是非同源的训练集,并有总共144个接口的化学位移扰动确定的残基。cons-PPISP预测了174个界面残基,准确率为69%,覆盖率为47%,并有望补充表征蛋白质-蛋白质界面的实验技术。
The number of structures of protein-protein complexes deposited to the Protein Data Bank is growing rapidly. These structures embed important information for predicting structures of new protein complexes. This motivated us to develop the PPISP method for predicting interface residues in protein-protein complexes. In PPISP, sequence profiles and solvent accessibility of spatially neighboring surface residues were used as input to a neural network. The network was trained on native interface residues collected from the Protein Data Bank. The prediction accuracy at the time was 70% with 47% coverage of native interface residues. Now we have extensively improved PPISP. The training set now consisted of 1156 nonhomologous protein chains. Test on a set of 100 nonhomologous protein chains showed that the prediction accuracy is now increased to 80% with 51% coverage. To solve the problem of over-prediction and under-prediction associated with individual neural network models, we developed a consensus method that combines predictions from multiple models with different levels of accuracy and coverage. Applied on a benchmark set of 68 proteins for protein protein docking, the consensus approach outperformed the best individual models by 3-8 percentage points in accuracy. To demonstrate the predictive power of cons-PPISP, eight complex-forming proteins with interfaces characterized by NMR were tested. These proteins are nonhomologous to the training set and have a total of 144 interface residues identified by chemical shift perturbation. cons-PPISP predicted 174 interface residues with 69% accuracy and 47% coverage and promises to complement experimental techniques in characterizing protein-protein interfaces.