Seeing the trees through the forest: sequence-based homo- and heteromeric protein-protein interaction sites prediction using random forest

Seeing the trees through the forest: sequence-based homo- and heteromeric protein-protein interaction sites prediction using random forest
复制标题

DOI:
10.1093/bioinformatics/btx005
复制
发表时间:
2017-05-15
期刊:
影响因子:
5.8
通讯作者:
Feenstra, K. Anton
Feenstra, K. Anton
中科院分区:
生物学3区
文献类型:
--
作者:
Hou, Qingzhen;De Geest, Paul F. G.;Feenstra, K. Anton

文献摘要

被引文献

相似文献

动机:基因组测序正在产生越来越多的相关蛋白质序列。然而,这些序列中很少有经过实验验证的注释,并且计算预测在生成此类注释方面变得越来越成功。一个关键的挑战仍然是预测给定蛋白质序列中参与蛋白质-蛋白质相互作用的氨基酸。此类预测通常基于机器学习方法,该方法利用已知参与相互作用的氨基酸的特性和序列位置。在本文中,我们使用随机森林(RF)评估各种特征的重要性,并将从序列预测的灵活性作为新的特征主干,以进一步优化蛋白质界面预测。结果:我们观察到,没有单一序列特征可以在我们的随机森林模型中精确定位相互作用位点。然而,结合不同的属性确实可以提高界面预测的性能。我们经过同源训练的 RF 界面预测器能够区分界面和非界面残基,在同源测试集中 ROC 曲线下面积为 0.72。经过异聚体训练的 RF 接口预测器在独立异聚体测试集上的表现优于现有预测器。我们在同聚体和异聚体组合数据集上训练了一个更通用的预测器,结果表明,除了预测同聚体界面之外,它还能够精确定位异二聚体中的界面残基。这表明我们的随机森林模型和特征包括捕获同二聚体和异二聚体界面的共同特性。补充信息:补充数据可在生物信息学在线获得。
Motivation: Genome sequencing is producing an ever-increasing amount of associated protein sequences. Few of these sequences have experimentally validated annotations, however, and computational predictions are becoming increasingly successful in producing such annotations. One key challenge remains the prediction of the amino acids in a given protein sequence that are involved in protein-protein interactions. Such predictions are typically based on machine learning methods that take advantage of the properties and sequence positions of amino acids that are known to be involved in interaction. In this paper, we evaluate the importance of various features using Random Forest (RF), and include as a novel feature backbone flexibility predicted from sequences to further optimise protein interface prediction.Results: We observe that there is no single sequence feature that enables pinpointing interacting sites in our Random Forest models. However, combining different properties does increase the performance of interface prediction. Our homomeric-trained RF interface predictor is able to distinguish interface from non-interface residues with an area under the ROC curve of 0.72 in a homomeric test-set. The heteromeric-trained RF interface predictor performs better than existing predictors on a independent heteromeric test-set. We trained a more general predictor on the combined homomeric and heteromeric dataset, and show that in addition to predicting homomeric interfaces, it is also able to pinpoint interface residues in heterodimers. This suggests that our random forest model and the features included capture common properties of both homodimer and heterodimer interfaces.Supplementary information: Supplementary data are available at Bioinformatics online.