Object-Oriented Implementation of?Algebraic Multi-grid Solver for Lattice QCD on SIMD Architectures and GPU Clusters

Object-Oriented Implementation of?Algebraic Multi-grid Solver for Lattice QCD on SIMD Architectures and GPU Clusters
复制标题

SIMD 架构和 GPU 集群上格子 QCD 代数多重网格求解器的面向对象实现

DOI:
10.1007/978-3-030-86976-2_15
复制
发表时间:
2021
期刊:
Lecture Notes in Computer Science
影响因子:
--
通讯作者:
Matsufuru Hideo
Matsufuru Hideo
中科院分区:
--
文献类型:
--
作者:
Kanamori Issaku;Ishikawa Ken-Ichi;Matsufuru Hideo

文献摘要

相似文献

精细算法的可移植实现对于在 HPC 应用程序中使用各种架构非常重要。在这项工作中,我们在三种不同的架构(Intel Xeon Phi、Fujitsu A64FX 和 NVIDIA Tesla V100)上实现了 Lattice QCD 的代数多网格求解器并进行了基准测试,以保持基于面向对象范例的代码的高性能和可移植性。代码的某些部分特定于采用适当的数据布局和调整的矩阵向量乘法内核的体系结构,而抽象求解器算法的实现对所有体系结构都是通用的。尽管求解器的性能取决于架构相关部分的调整,但我们观察到合理的缩放行为和比混合精度 BiCGSstab 求解器更好的性能。
A portable implementation of elaborated algorithm is important to use variety of architectures in HPC applications. In this work we implement and benchmark an algebraic multi-grid solver for Lattice QCD on three different architectures, Intel Xeon Phi, Fujitsu A64FX, and NVIDIA Tesla V100, in keeping high performance and portability of the code based on the object-oriented paradigm. Some parts of code are specific to an architecture employing appropriate data layout and tuned matrix-vector multiplication kernels, while the implementation of abstract solver algorithm is common to all architectures. Although the performance of the solver depends on tuning of the architecture-dependent part, we observe reasonable scaling behavior and better performance than the mixed precision BiCGSstab solvers.