Improving prediction of heterodimeric protein complexes using combination with pairwise kernel.

Improving prediction of heterodimeric protein complexes using combination with pairwise kernel.
复制标题

DOI:
10.1186/s12859-018-2017-5
复制
发表时间:
2018-02-19
期刊:
影响因子:
3
通讯作者:
Vert JP
Vert JP
中科院分区:
生物学4区
文献类型:
--
作者:
Ruan P;Hayashida M;Akutsu T;Vert JP

文献摘要

参考文献

被引文献

相似文献

由于许多蛋白质只有在与其伴侣蛋白相互作用并形成蛋白质复合物后才具有功能,因此识别形成复合物的蛋白质组至关重要。因此,已经提出了几种计算方法来预测复合物的实验蛋白质-蛋白质相互作用(PPI)网络的拓扑结构和结构。这些方法很好地预测涉及至少三种蛋白质的复合物,但通常无法识别仅涉及两种不同蛋白质的复合物,称为异二聚体复合物或异二聚体。然而,迫切需要有效的方法来预测异二聚体,因为大多数已知的蛋白质复合物正是异二聚体。在本文中,我们使用了三个有前途的核函数,最小核和两个成对核,这是度量学习成对核(MLPK)和张量积成对核(TPPK)。我们还考虑了Min核的归一化形式。然后,通过插补将联合收割机Min核或其归一化形式与其中一个成对核组合。我们应用核PPI的基础上,域,系统发育概况,和亚细胞定位特性来预测异源二聚体。然后,我们评估我们的方法,采用C-支持向量分类(C-SVC),进行10倍交叉验证,并计算平均F-措施。结果表明,归一化最小核和MLPK的组合导致最好的F-措施,并提高了我们以前的工作,这是迄今为止最好的现有方法的性能。我们提出了新的方法来预测异源二聚体,使用基于机器学习的方法。我们训练了一个支持向量机(SVM)来区分相互作用与非相互作用的蛋白质对,基于从PPI,结构域,系统发育概况和亚细胞定位提取的信息。我们详细评估了新的内核函数来编码这些数据,并报告了优于最先进技术的预测性能。
Since many proteins become functional only after they interact with their partner proteins and form protein complexes, it is essential to identify the sets of proteins that form complexes. Therefore, several computational methods have been proposed to predict complexes from the topology and structure of experimental protein-protein interaction (PPI) network. These methods work well to predict complexes involving at least three proteins, but generally fail at identifying complexes involving only two different proteins, called heterodimeric complexes or heterodimers. There is however an urgent need for efficient methods to predict heterodimers, since the majority of known protein complexes are precisely heterodimers. In this paper, we use three promising kernel functions, Min kernel and two pairwise kernels, which are Metric Learning Pairwise Kernel (MLPK) and Tensor Product Pairwise Kernel (TPPK). We also consider the normalization forms of Min kernel. Then, we combine Min kernel or its normalization form and one of the pairwise kernels by plugging. We applied kernels based on PPI, domain, phylogenetic profile, and subcellular localization properties to predicting heterodimers. Then, we evaluate our method by employing C-Support Vector Classification (C-SVC), carrying out 10-fold cross-validation, and calculating the average F-measures. The results suggest that the combination of normalized-Min-kernel and MLPK leads to the best F-measure and improved the performance of our previous work, which had been the best existing method so far. We propose new methods to predict heterodimers, using a machine learning-based approach. We train a support vector machine (SVM) to discriminate interacting vs non-interacting protein pairs, based on informations extracted from PPI, domain, phylogenetic profiles and subcellular localization. We evaluate in detail new kernel functions to encode these data, and report prediction performance that outperforms the state-of-the-art.
DOI: 10.1038/nature04532
发表时间: 2006-03-30
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Aloy, P;Superti-Furga, G
通讯作者: Superti-Furga, G
DOI: 10.1142/s0219720008003497
发表时间: 2008-06-01
影响因子: 1
作者:
Chua, Hon Nian;Ning, Kang;Wong, Limsoon
通讯作者: Wong, Limsoon
DOI: 10.1186/1471-2105-4-2
发表时间: 2003-01-13
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Bader, GD;Hogue, CW
通讯作者: Hogue, CW
DOI: 10.1073/pnas.96.8.4285
发表时间: 1999-04-13
影响因子: 11.1
作者:
Pellegrini, M;Marcotte, EM;Yeates, TO
通讯作者: Yeates, TO
DOI: 10.1186/1477-5956-9-s1-s14
发表时间: 2011-10-14
期刊: Proteome science
影响因子: 2
作者:
Maruyama O;Chihara A
通讯作者: Chihara A