MPI + OpenACC: Accelerating radiation transport mini-application, minisweep, on heterogeneous systems

MPI + OpenACC: Accelerating radiation transport mini-application, minisweep, on heterogeneous systems
复制标题

MPI OpenACC:加速异构系统上的辐射传输微型应用程序、minisweep

DOI:
10.1016/j.cpc.2018.10.007
复制
发表时间:
2019
影响因子:
6.3
通讯作者:
Hernandez, Oscar
Hernandez, Oscar
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
Searles, Robert;Chandrasekaran, Sunita;Joubert, Wayne;Hernandez, Oscar

文献摘要

参考文献

被引文献

相似文献

体系结构正在迅速发展,预计十亿级机器将提供十亿路并发。我们需要重新考虑算法、语言和编程模型等组件,以便迁移大规模应用程序并探索这些机器上的并行性。虽然基于指令的编程模型让程序员更少地关注编程,更多地关注科学,但在这些模型中表达复杂的并行模式可能是一项艰巨的任务,特别是当目标是匹配硬件平台可以提供的性能时。一种这样的模式是波前。本文深入研究了Denovo的一个基于波前的微型应用程序,Denovo是一个用于核反应堆建模的生产代码。我们使用CUDA 9.0,OpenMP 4.0(SIMD)和OpenACC 2.6在Minisweep(miniapplication)的主内核中并行化Koch-Baker-Alcouffe(KBA)并行波前扫描算法。我们在NVIDIA下一代Volta GPU上运行的OpenACC实现在串行代码上的加速比为85.06倍,高于CUDA在相同串行实现上的83.72倍。我们还探索了我们的解决方案的可扩展性,使用MPI来分解我们的仿真域,使我们能够在最先进的HPC系统中的许多节点和加速器上运行。我们跨平台的并行化工作也促使我们定义一个抽象的并行模型,是架构独立的,创建软件抽象,可以使用的应用程序采用波前扫描图案的目标。程序摘要程序标题:Minisweep程序文件doi:http://dx。doi。org/10.17632/cbcp37t8gf。1许可证条款:BSD 2-clause编程语言:C问题性质:Minisweep代理应用程序[1]是Profugus辐射传输miniapp项目[2]的一部分,该项目再现了Denovo S n辐射传输代码的扫描内核的计算模式[3]。扫描内核负责Denovo的大部分计算开销(80%-99%)。Denovo是核反应堆中子学建模的生产代码,目前正在由DOE INCITE项目用于对国际热核实验反应堆(ITER)聚变反应堆进行建模[4]。在高节点数下执行反应堆模拟所需的此代码的多次运行使其成为高效映射到加速架构的重要目标。解决方法:这项工作提出了一个抽象的并行模型,有效地映射波前应用程序到现代HPC架构。Minisweep被用作评估这种技术的案例研究。我们的评估是使用OpenACC来针对许多架构进行的。
Architectures are rapidly evolving, and exascale machines are expected to offer billion-way concurrency. We need to rethink algorithms, languages and programming models among other components in order to migrate large scale applications and explore parallelism on these machines. Although directive-based programming models allow programmers to worry less about programming and more about science, expressing complex parallel patterns in these models can be a daunting task especially when the goal is to match the performance that the hardware platforms can offer. One such pattern is wavefront. This paper extensively studies a wavefront-based miniapplication for Denovo, a production code for nuclear reactor modeling. We parallelize the Koch–Baker–Alcouffe (KBA) parallel-wavefront sweep algorithm in the main kernel of Minisweep (the miniapplication) using CUDA 9.0, OpenMP 4.0 (SIMD) and OpenACC 2.6. Our OpenACC implementation running on NVIDIA’s next-generation Volta GPU boasts an 85.06 x speedup over serial code, which is larger than CUDA’s 83.72 x speedup over the same serial implementation. We also explore the scalability of our solution using MPI to decompose our simulation domain, allowing us to run on many nodes and accelerators present in state-of-the-art HPC systems. Our parallelization effort across platforms also motivated us to define an abstract parallelism model that is architecture independent, with a goal of creating software abstractions that can be used by applications employing the wavefront sweep motif. Program summary Program Title: Minisweep Program Files doi: http://dx. doi. org/10.17632/cbcp37t8gf. 1 Licensing provisions: BSD 2-clause Programming language: C Nature of problem: The Minisweep proxy application [1] is part of the Profugus radiation transport miniapp project [2] that reproduces the computational pattern of the sweep kernel of the Denovo S n radiation transport code [3]. The sweep kernel is responsible for most of the computational expense (80%–99%) of Denovo. Denovo, a production code for nuclear reactor neutronics modeling, is in use by a current DOE INCITE project to model the International Thermonuclear Experimental Reactor (ITER) fusion reactor [4]. The many runs of this code required to perform reactor simulations at high node counts make it an important target for efficient mapping to accelerated architectures. Solution method: This work proposes an abstract parallelism model for efficiently mapping wavefront application to modern HPC architectures. Minisweep is used as a case study for evaluating this technique. Our evaluation is performed using OpenACC to target many architectures.
分布式内存多处理器上三角系统的并行求解
DOI: --
发表时间: 1988
期刊:
影响因子: --
作者:
M. Heath;C. Romine
通讯作者: C. Romine
Cell 宽带引擎上的并行 DNA 序列比对
DOI: --
发表时间: 2007
期刊: Parallel Processing and Applied Mathematics
影响因子: --
作者:
A. Wirawan;C. Kwoh;B. Schmidt
通讯作者: B. Schmidt
C2FPGA - 依赖时序图设计方法
DOI: --
发表时间: 2013
期刊: J. Parallel Distributed Comput.
影响因子: --
作者:
S. Chandrasekaran;Shilpa Shanbagh;R. Jayaraman;D. Maskell;Hui Yan Cheah
通讯作者: Hui Yan Cheah
DOI: --
发表时间: 2013
期刊:
影响因子: --
作者:
R. Baker
通讯作者: R. Baker
Swift/T:通过分布式内存数据流处理构建大规模应用程序
DOI: 10.1109/ccgrid.2013.99
发表时间: 2013
期刊: 2013 13th IEEE/ACM International Symposium on Cluster, Cloud, and Grid Computing
影响因子: --
作者:
J. Wozniak;Timothy G. Armstrong;M. Wilde;D. Katz;E. Lusk;Ian T Foster
通讯作者: Ian T Foster