ReSAPP: Predicting overlapping protein complexes by merging multiple-sampled partitions of proteins

ReSAPP: Predicting overlapping protein complexes by merging multiple-sampled partitions of proteins
复制标题

DOI:
10.1142/s0219720014420049
复制
发表时间:
2014-12-01
影响因子:
1
通讯作者:
Maruyama, Osamu
Maruyama, Osamu
中科院分区:
生物学4区
文献类型:
--
作者:
Kobiki, So;Maruyama, Osamu

文献摘要

被引文献

相似文献

已知许多蛋白质在形成特定的蛋白质组(称为蛋白质复合物)时执行其自身的功能。随着大规模蛋白质-蛋白质相互作用(PPI)研究的出现,从PPI预测蛋白质复合物已成为系统生物学中一个具有挑战性的问题。在本文中,我们提出了一种新的方法,称为重复模拟退火的蛋白质分区(ReSAPP),预测蛋白质复合物从加权PPI。在第一阶段,ReSAPP通过重复地将基于模拟退火的优化算法应用于PPI来生成给定PPI的所有蛋白质的多个(可能不同的)分区。在第二阶段,将这些多个分区中大小为二或更多的所有不同集群合并到这些集群的集合中,这些集群作为预测的蛋白质复合物输出。在ReSAPP与我们以前的算法PPSampler 2以及其他各种工具MCL,MCODE,DPClus,CMC,COACH,RRW,NWE和PPSampler 1的性能比较中,ReSAPP优于其他方法。此外,ReSAPP的F-度量值高于不合并分区的ReSAPP变体的F-度量值。因此,我们凭经验得出结论,采样多个分区并合并它们的组合对于预测蛋白质复合物是有效的。
Many proteins are known to perform their own functions when they form particular groups of proteins, called protein complexes. With the advent of large-scale protein-protein interaction (PPI) studies, it has been a challenging problem in systems biology to predict protein complexes from PPIs. In this paper, we propose a novel method, called Repeated Simulated Annealing of Partitions of Proteins (ReSAPP), which predicts protein complexes from weighted PPIs. ReSAPP, in the first stage, generates multiple (possibly different) partitions of all proteins of given PPIs by repeatedly applying a simulated annealing based optimization algorithm to the PPIs. In the second stage, all different clusters of size two or more in those multiple partitions are merged into a collection of those clusters, which are outputted as predicted protein complexes. In performance comparison of ReSAPP with our previous algorithm, PPSampler2, as well as other various tools, MCL, MCODE, DPClus, CMC, COACH, RRW, NWE, and PPSampler1, ReSAPP is shown to outperform the other methods. Furthermore, the value of F-measure of ReSAPP is higher than that of the variant of ReSAPP without merging partitions. Thus, we empirically conclude that the combination of sampling multiple partitions and merging them is effective to predict protein complexes.