Using Advanced Vector Extensions AVX-512 for MPI Reductions

Using Advanced Vector Extensions AVX-512 for MPI Reductions
复制标题

使用高级矢量扩展 AVX-512 减少 MPI

DOI:
10.1145/3416315.3416316
复制
发表时间:
2020
期刊:
EuroMPI/USA '20: 27th European MPI Users' Group Meeting
影响因子:
--
通讯作者:
Dongarra, Jack
Dongarra, Jack
中科院分区:
--
文献类型:
--
作者:
Zhong, Dong;Cao, Qinglei;Bosilca, George;Dongarra, Jack

文献摘要

参考文献

被引文献

相似文献

随着高性能计算(HPC)系统规模的不断增长,研究人员致力于探索提高并行性水平以实现最佳性能。现代CPU的设计,包括其分层存储和SIMD/矢量化能力的特点,决定了算法的效率。最近引入的宽向量指令集扩展(AVX和SVE)促使向量化成为提高效率和缩小与峰值性能差距的关键。在本文中,我们提出了一种预定义的MPI约简操作的实现,利用AVX, AVX2和AVX-512本征来提供基于向量的约简操作,并改善这些预定义的MPI约简操作的求解时间。通过这些优化,我们实现了更高的本地计算效率,这直接有利于集体减少的总体成本。对不同场景下的软件栈进行了评估,结果表明该方案具有通用性和高效性。在Intel至强Gold集群上进行的实验表明,AVX-512优化的缩减操作比Open MPI默认的MPI本地缩减实现了10倍的性能优势。
As the scale of high-performance computing (HPC) systems continues to grow, researchers are devoted themselves to explore increasing levels of parallelism to achieve optimal performance. The modern CPU’s design, including its features of hierarchical memory and SIMD/vectorization capability, governs algorithms’ efficiency. The recent introduction of wide vector instruction set extensions (AVX and SVE) motivated vectorization to become of critical importance to increase efficiency and close the gap to peak performance.In this paper, we propose an implementation of predefined MPI reduction operations utilizing AVX, AVX2 and AVX-512 intrinsics to provide vector-based reduction operation and to improve the time-to-solution of these predefined MPI reduction operations. With these optimizations, we achieve higher efficiency for local computations, which directly benefit the overall cost of collective reductions. The evaluation of the resulting software stack under different scenarios demonstrates that the solution is at the same time generic and efficient. Experiments are conducted on an Intel Xeon Gold cluster, which shows our AVX-512 optimized reduction operations achieve 10X performance benefits than Open MPI default for MPI local reduction.
在模板代码上使用 Arm 的可扩展矢量扩展
DOI: --
发表时间: 2019
影响因子: 3.3
作者:
Adrià Armejach;Helena Caminal;J. M. Cebrian;Rubén Langarita;Rekai González;Chris Adeniyi;M. Valero;Marc Casas;Miquel Moretó
通讯作者: Miquel Moretó
DOI: 10.1145/2837614.2837615
发表时间: 2016-01
期刊: Proceedings of the 43rd Annual ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages
影响因子: --
作者:
Shaked Flur;Kathryn E. Gray;Christopher Pulte;Susmit Sarkar;A. Sezgin;Luc Maranget;Will Deacon;Peter Sewell
通讯作者: Shaked Flur;Kathryn E. Gray;Christopher Pulte;Susmit Sarkar;A. Sezgin;Luc Maranget;Will Deacon;Peter Sewell
矢量长度不可知架构上的模板代码
DOI: --
发表时间: 2018
期刊: International Conference on Parallel Architectures and Compilation Techniques
影响因子: --
作者:
Adrià Armejach;Helena Caminal;J. M. Cebrian;Rekai González;Chris Adeniyi;M. Valero;Marc Casas;Miquel Moretó
通讯作者: Miquel Moretó
自动向量化编译器的比较研究
DOI: 10.1016/s0167-8191(05)80035-3
发表时间: 1991
期刊: Parallel Comput.
影响因子: --
作者:
D. Levine;D. Callahan;J. Dongarra
通讯作者: J. Dongarra
稀疏浮点数据的 MPI 约简运算
DOI: --
发表时间: 2008
期刊: PVM/MPI
影响因子: --
作者:
Michael Hofmann;G. Rünger
通讯作者: G. Rünger