Fuse: multiple network alignment via data fusion

Fuse: multiple network alignment via data fusion
复制标题

DOI:
10.1093/bioinformatics/btv731
复制
发表时间:
2016-04-15
期刊:
影响因子:
5.8
通讯作者:
Przulj, Natasa
Przulj, Natasa
中科院分区:
生物学3区
文献类型:
--
作者:
Gligorijevic, Vladimir;Malod-Dognin, Noel;Przulj, Natasa

文献摘要

被引文献

相似文献

动机:发现蛋白质-蛋白质相互作用(PPI)网络中的模式是系统生物学的核心问题。这些网络之间的比对有助于功能理解,因为它们揭示了重要的信息,如进化保守途径,蛋白质复合物和功能直系同源物。然而,多网络比对问题的复杂性随着被比对的网络的数量呈指数增长,并且设计既可扩展又产生生物学相关比对的多网络比对器是一项尚未完全解决的具有挑战性的任务。多网络对齐的目标是创建在所有网络中进化和功能保守的节点集群。不幸的是,迄今为止提出的对齐方法不符合这一目标,因为它们是由成对的分数,不利用整个功能和进化信息在所有networks.Results:为了克服这一弱点,我们提出了一个新的多网络对齐算法,工作在两个步骤。首先,它通过融合来自所有对齐的PPI网络的布线模式及其蛋白质之间的序列相似性的信息来计算我们的新蛋白质功能相似性得分。这与之前的工具相反,这些工具都是基于正在对齐的网络对中的蛋白质相似性。我们全面的新蛋白质相似性得分是通过非负矩阵三因子分解(NMTF)方法计算的,该方法预测蛋白质之间的关联,这些蛋白质的同源性(来自序列)和功能相似性(来自布线模式)得到所有网络的支持。使用BioGRID的五个最大和最完整的PPI网络,我们表明NMTF预测了大量生物学上一致的蛋白质对。其次,为了在所有网络中识别比对蛋白质的簇,我们使用了我们新的最大权重k-部匹配近似算法。我们比较了多个网络比对器与最先进的多个网络比对器,并表明(i)通过仅使用序列比对评分,多个网络比对器已经优于其他比对器,并产生了大量的生物学上一致的聚类,覆盖了所有对齐的PPI网络,(ii)使用序列比对和拓扑NMTF预测的评分,导致迄今为止最好的多个网络比对。
Motivation: Discovering patterns in networks of protein-protein interactions (PPIs) is a central problem in systems biology. Alignments between these networks aid functional understanding as they uncover important information, such as evolutionary conserved pathways, protein complexes and functional orthologs. However, the complexity of the multiple network alignment problem grows exponentially with the number of networks being aligned and designing a multiple network aligner that is both scalable and that produces biologically relevant alignments is a challenging task that has not been fully addressed. The objective of multiple network alignment is to create clusters of nodes that are evolutionarily and functionally conserved across all networks. Unfortunately, the alignment methods proposed thus far do not meet this objective as they are guided by pairwise scores that do not utilize the entire functional and evolutionary information across all networks.Results: To overcome this weakness, we propose Fuse, a new multiple network alignment algorithm that works in two steps. First, it computes our novel protein functional similarity scores by fusing information from wiring patterns of all aligned PPI networks and sequence similarities between their proteins. This is in contrast with the previous tools that are all based on protein similarities in pairs of networks being aligned. Our comprehensive new protein similarity scores are computed by Non-negative Matrix Tri-Factorization (NMTF) method that predicts associations between proteins whose homology (from sequences) and functioning similarity (from wiring patterns) are supported by all networks. Using the five largest and most complete PPI networks from BioGRID, we show that NMTF predicts a large number protein pairs that are biologically consistent. Second, to identify clusters of aligned proteins over all networks, Fuse uses our novel maximum weight k-partite matching approximation algorithm. We compare Fuse with the state of the art multiple network aligners and show that (i) by using only sequence alignment scores, Fuse already outperforms other aligners and produces a larger number of biologically consistent clusters that cover all aligned PPI networks and (ii) using both sequence alignments and topological NMTF-predicted scores leads to the best multiple network alignments thus far.