Pencil: A Pipelined Algorithm for Distributed Stencils

Pencil: A Pipelined Algorithm for Distributed Stencils
复制标题

DOI:
10.1109/sc41405.2020.00089
复制
发表时间:
2020-11
期刊:
SC20: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Hengjie Wang;Aparna Chandramowlishwaran
Hengjie Wang;Aparna Chandramowlishwaran
中科院分区:
其他
文献类型:
--
作者:
Hengjie Wang;Aparna Chandramowlishwaran

文献摘要

被引文献

相似文献

Stenointment计算是各种计算流体动力学(CFD)应用的核心,并且已经被充分研究了几十年。通常,它们是高度内存限制的,因此,已经提出了许多平铺算法来提高其性能。尽管高效,但这些算法中的大多数都是为共享内存机器上的单次迭代空间而设计的。然而,在计算流体力学中,我们面临着多块结构的网格组成的多个连接的迭代空间分布在许多节点,本文中,我们提出了一种流水线模板算法称为分布式存储机,适用于实际的计算流体力学问题,跨越多个迭代空间。基于对单个节点上的高速缓存平铺的深入分析,我们首先确定了MPI和OpenMP用于时间平铺的最佳组合和最佳平铺方法,其性能优于最先进的自动并行化工具Pluto,最高可达1.92\times$。然后,我们采用DeepHalo来解耦多个连接的迭代空间,以便可以将时间平铺应用于每个空间。最后,我们通过流水线的计算和通信实现重叠,而不牺牲时间缓存平铺的优势。在两台具有Omni-Path和InfiniBand网络的分布式内存机器上,使用4种跨6种数值方案的stenosis来评估网络性能。在Omni-Path系统上,对于多达128个节点,OpenGL表现出出色的弱可扩展性和强可扩展性,并且在具有32个节点的多块网格上比具有空间平铺的MPI+OpenMP Funneled高出1.33 - 3.41 \times$。
Stencil computations are at the core of various Computational Fluid Dynamics (CFD) applications and have been well-studied for several decades. Typically they’re highly memory-bound and as a result, numerous tiling algorithms have been proposed to improve its performance. Although efficient, most of these algorithms are designed for single iteration spaces on shared-memory machines. However, in CFD, we are confronted with multi-block structured girds composed of multiple connected iteration spaces distributed across many nodes.In this paper, we propose a pipelined stencil algorithm called Pencil for distributed memory machines that applies to practical CFD problems that span multiple iteration spaces. Based on an in-depth analysis of cache tiling on a single node, we first identify both the optimal combination of MPI and OpenMP for temporal tiling and the best tiling approach, which outperforms the state-of-the-art automatic parallelization tool Pluto by up to $1.92 \times$. Then, we adopt DeepHalo to decouple the multiple connected iteration spaces so that temporal tiling can be applied to each space. Finally, we achieve overlap by pipelining the computation and communication without sacrificing the advantage from temporal cache tiling. Pencil is evaluated using 4 stencils across 6 numerical schemes on two distributed memory machines with Omni-Path and InfiniBand networks. On the Omni-Path system, Pencil exhibits outstanding weak and strong scalability for up to 128 nodes and outperforms MPI+OpenMP Funneled with space tiling by $1.33- 3.41 \times$ on a multi-block grid with 32 nodes.