HipMCL: a high-performance parallel implementation of the Markov clustering algorithm for large-scale networks.

HipMCL: a high-performance parallel implementation of the Markov clustering algorithm for large-scale networks.
复制标题

DOI:
10.1093/nar/gkx1313
复制
发表时间:
2018-04-06
影响因子:
14.9
通讯作者:
Buluç A
Buluç A
中科院分区:
生物学2区
文献类型:
--
作者:
Azad A;Pavlopoulos GA;Ouzounis CA;Kyrpides NC;Buluç A

文献摘要

参考文献

被引文献

相似文献

生物网络捕获相关实体(例如分子、蛋白质或基因)的结构或功能特性。典型的例子是基因表达网络或蛋白质-蛋白质相互作用网络,它们保存有关功能亲和力或结构相似性的信息。由于生物数据的规模和丰富性不断增加,此类网络的规模不断扩大。虽然已经提出了各种聚类算法来查找高度连接的区域,但马尔可夫聚类 (MCL) 一直是对序列相似性或表达网络进行聚类的最成功的方法之一。尽管很受欢迎,但由于运行时间长和内存需求高,MCL 对大型数据集进行集群的可扩展性仍然是一个瓶颈。在这里,我们提出了高性能 MCL (HipMCL),它是原始 MCL 算法的并行实现,可以在分布式内存计算机上运行。我们证明 HipMCL 可以在 2.4 小时内有效地利用 2000 个计算节点,并聚集一个由 7000 万个节点和 680 亿条边组成的网络。通过利用分布式内存环境,HipMCL 集群大型网络的速度比 MCL 快几个数量级,并且能够集群更大的网络。 HipMCL 基于 MPI 和 OpenMP,可在修改后的 BSD 许可证下免费使用。
Biological networks capture structural or functional properties of relevant entities such as molecules, proteins or genes. Characteristic examples are gene expression networks or protein–protein interaction networks, which hold information about functional affinities or structural similarities. Such networks have been expanding in size due to increasing scale and abundance of biological data. While various clustering algorithms have been proposed to find highly connected regions, Markov Clustering (MCL) has been one of the most successful approaches to cluster sequence similarity or expression networks. Despite its popularity, MCL’s scalability to cluster large datasets still remains a bottleneck due to high running times and memory demands. Here, we present High-performance MCL (HipMCL), a parallel implementation of the original MCL algorithm that can run on distributed-memory computers. We show that HipMCL can efficiently utilize 2000 compute nodes and cluster a network of ∼70 million nodes with ∼68 billion edges in ∼2.4 h. By exploiting distributed-memory environments, HipMCL clusters large-scale networks several orders of magnitude faster than MCL and enables clustering of even bigger networks. HipMCL is based on MPI and OpenMP and is freely available under a modified BSD license.
DOI: 10.1101/gr.113985.110
发表时间: 2011-03-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Kielbasa, Szymon M.;Wan, Raymond;Frith, Martin C.
通讯作者: Frith, Martin C.
DOI: 10.1177/1094342011403516
发表时间: 2011-11-01
影响因子: 3.1
作者:
Buluc, Aydin;Gilbert, John R.
通讯作者: Gilbert, John R.
DOI: 10.1093/nar/gkw929
发表时间: 2017-01-04
影响因子: 14.9
作者:
Chen IA;Markowitz VM;Chu K;Palaniappan K;Szeto E;Pillay M;Ratner A;Huang J;Andersen E;Huntemann M;Varghese N;Hadjithomas M;Tennessen K;Nielsen T;Ivanova NN;Kyrpides NC
通讯作者: Kyrpides NC
DOI: 10.1186/1471-2105-4-2
发表时间: 2003-01-13
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Bader, GD;Hogue, CW
通讯作者: Hogue, CW
DOI: 10.1088/1742-5468/2008/10/p10008
发表时间: 2008-10-01
影响因子: 2.4
作者:
Blondel, Vincent D.;Guillaume, Jean-Loup;Lefebvre, Etienne
通讯作者: Lefebvre, Etienne