Evaluation of a Minimally Synchronous Algorithm for 2:1 Octree Balance

Evaluation of a Minimally Synchronous Algorithm for 2:1 Octree Balance
复制标题

2:1 八叉树平衡的最小同步算法的评估

DOI:
--
复制
发表时间:
2020
期刊:
International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
T. Isaac
T. Isaac
中科院分区:
--
文献类型:
--
作者:
Hansol Suh;T. Isaac

文献摘要

被引文献

相似文献

p4est库实现了基于八叉树的自适应网格细化(AMR),并在以前的弱缩放研究中证明了超过100,000个MPI进程的并行可扩展性。这项工作的重点是强大的可扩展性的网格自适应p4est,其中现有的2:1平衡的通信模式是一个延迟瓶颈。基于排序的Malhotra和Biros算法具有通信均衡性,但进程间存在冲突。我们提出了一个结合排序和邻居交换的算法,以最小化每个进程所占用的进程数,并在TACC的Stampede2上测试了这些算法的性能。并行排序和最小同步算法的性能都明显优于现有算法,并且在1,024个Xeon Phi KNL节点上具有几乎相同的性能,这意味着最小同步算法的渐进优势不会转化为这种规模的性能改进。我们的结论是,全球元数据通信将限制未来的强大的伸缩性。
The p4est library implements octree-based adaptive mesh refinement (AMR) and has demonstrated parallel scalability beyond 100,000 MPI processes in previous weak scaling studies. This work focuses on the strong scalability of mesh adaptivity in p4est, where the communication pattern of the existing 2:1-balance is a latency bottleneck. The sorting-based algorithm of Malhotra and Biros has balanced communication, but synchronizes all processes. We propose an algorithm that combines sorting and neighbor-to-neighbor exchange to minimize the number of processes each process synchronizes with.We measure the performance of these algorithms on several test problems on Stampede2 at TACC. Both the parallel-sorting and minimally-synchronous algorithms significantly outperform the existing algorithm and have nearly identical performance out to 1,024 Xeon Phi KNL nodes, meaning the asymptotic advantage of the minimally-synchronous algorithm does not translate to improved performance at this scale. We conclude by showing that global metadata communication will limit future strong scaling.