A DISTRIBUTED-MEMORY ALGORITHM FOR COMPUTING A HEAVY-WEIGHT PERFECT MATCHING ON BIPARTITE GRAPHS

A DISTRIBUTED-MEMORY ALGORITHM FOR COMPUTING A HEAVY-WEIGHT PERFECT MATCHING ON BIPARTITE GRAPHS
复制标题

DOI:
10.1137/18m1189348
复制
发表时间:
2020-01-01
影响因子:
3.1
通讯作者:
Langguth, Johannes
Langguth, Johannes
中科院分区:
数学2区
文献类型:
--
作者:
Azad, Ariful;Buluc, Aydin;Langguth, Johannes

文献摘要

被引文献

相似文献

我们设计并实现了一种在加权二部图中寻找完美匹配的高效并行算法,使得匹配的边上的权重很大。这个问题不同于最大权匹配问题,后者的可伸缩近似算法是已知的。它的主要动机是在因子分解之前在可伸缩的稀疏直接求解器中找到好的支点。由于缺乏可伸缩的替代方案,分布式解算器使用最大权重完美匹配算法的顺序实现,例如MC64中提供的那些算法。为了克服这一局限性,我们提出了一种完全并行的分布式存储算法,该算法首先生成完美匹配,然后通过并行搜索长度为4的增权循环来迭代地提高完美匹配的权重。对于大多数实际问题,我们的算法生成的完美匹配的权值都非常接近最优。该算法的高效实现在Cray XC40超级计算机上扩展到256个节点(17,408个核心),并且可以解决太大而无法使用顺序算法的单个节点处理的实例。
We design and implement an efficient parallel algorithm for finding a perfect matching in a weighted bipartite graph such that weights on the edges of the matching are large. This problem differs from the maximum weight matching problem, for which scalable approximation algorithms are known. It is primarily motivated by finding good pivots in scalable sparse direct solvers before factorization. Due to the lack of scalable alternatives, distributed solvers use sequential implementations of maximum weight perfect matching algorithms, such as those available in MC64. To overcome this limitation, we propose a fully parallel distributed memory algorithm that first generates a perfect matching and then iteratively improves the weight of the perfect matching by searching for weight-increasing cycles of length 4 in parallel. For most practical problems the weights of the perfect matchings generated by our algorithm are very close to the optimum. An efficient implementation of the algorithm scales up to 256 nodes (17,408 cores) on a Cray XC40 supercomputer and can solve instances that are too large to be handled by a single node using the sequential algorithm.