Whole Program Generation of Massively Parallel Shallow Water Equation Solvers

Whole Program Generation of Massively Parallel Shallow Water Equation Solvers
复制标题

大规模并行浅水方程求解器的整个程序生成

DOI:
--
复制
发表时间:
2018
期刊:
IEEE International Conference on Cluster Computing
影响因子:
--
通讯作者:
H. Köstler
H. Köstler
中科院分区:
--
文献类型:
--
作者:
S. Kuckuk;H. Köstler

文献摘要

参考文献

被引文献

相似文献

几十年来,洋流的研究一直是一个活跃的研究领域。浅水方程作为一种接近水面的模型,可以采用。对于真实的模拟,高效的数值求解器是必要的,表现出良好的节点级性能,同时仍然保持可扩展性。当比较离散化模型和实际实现时,人们经常会发现它们有很大的不同。这种差距使得领域专家很难实现他们的模型,需要高性能计算(HPC)专家来确保最佳的实现。使用领域特定语言(DSL)和代码生成技术是弥补这一差距的有用工具。近年来,ExaStengland及其DSL ExaSlang已被证明为此提供了合适的平台。我们提出了一个扩展到现在的椭圆双曲型偏微分方程(PDE)在这项工作中,即SWE。在设置了合适的离散化之后,我们演示了如何将其映射到ExaSlang代码。这些代码仍然与原始的、数学上的规范非常相似,并且可以由领域专家轻松编写。尽管如此,从这种抽象表示生成的求解器可以在大规模集群上运行。我们通过在最先进的GPU集群Piz Daint上提供性能和可扩展性结果来证明这一点,我们在2048个GPU上解决了近万亿个未知数。从那里,我们讨论了不同的优化,如重叠的计算和通信,或切换到一个混合的CPU-GPU并行化方案的性能影响。
The study of ocean currents has been an active area of research for decades. As a model close to the water surface, the shallow water equations (SWE) can be used. For realistic simulations, efficient numerical solvers are necessary that exhibit a good node-level performance while still maintaining scalability. When comparing the discretized model and the actual implementation, one often finds that they differ vastly. This gap makes it hard for domain experts to implement their models and high performance computing (HPC) experts are required to ensure an optimal implementation. Using domain-specific languages (DSLs) and code generation techniques can be a useful tool to bridge this gap. In recent years, ExaStencils and its DSL ExaSlang have proven to provide a suitable platform for this. We present an extension from up to now elliptic to hyperbolic partial differential equations (PDEs) in this work, namely the SWE. After setting up a suitable discretization, we demonstrate how it can be mapped to ExaSlang code. This code is still quite similar to the original, mathematically motivated specification and can be easily written by domain experts. Still, solvers generated from this abstract representation can be run on large-scale clusters. We demonstrate this by giving performance and scalability results on the state-of-the-art GPU cluster Piz Daint where we solve for close to a trillion unknowns on 2048 GPUs. From there, we discuss the performance impact of different optimizations such as overlapping computation and communication, or switching to a hybrid CPU-GPU parallelization scheme.
DOI: 10.1109/jproc.2018.2854229
发表时间: 2018
影响因子: 20.6
作者:
Schmitt;Kronawitter;Hannig;Lengauer
通讯作者: Lengauer