Evaluation of Directive-Based GPU Programming Models on a Block Eigensolver with Consideration of Large Sparse Matrices

Evaluation of Directive-Based GPU Programming Models on a Block Eigensolver with Consideration of Large Sparse Matrices
复制标题

考虑大型稀疏矩阵的块特征求解器上基于指令的 GPU 编程模型的评估

DOI:
10.1007/978-3-030-49943-3_4
复制
发表时间:
2020
期刊:
Lecture Notes in Computer Science: Accelerator Programming Using Directives (WACCPD 2019
影响因子:
--
通讯作者:
Wright, Nicholas J.
Wright, Nicholas J.
中科院分区:
--
文献类型:
--
作者:
Rabbi, Fazlay;Daley, Christopher S.;Aktulga, Hasan Metin;Wright, Nicholas J.

文献摘要

参考文献

被引文献

相似文献

实现大规模科学应用的高性能和性能可移植性是异构计算系统(例如多核CPU和GPU等加速器)面临的主要挑战。在这项工作中,我们使用两种流行的基于指令的编程模型(OpenMP 和 OpenACC)用于 GPU 加速系统,实现了一种广泛使用的块特征求解器,即局部最优块预条件共轭梯度 (LOBPCG)。我们的工作与现有工作的不同之处在于,它采用整体方法来优化整个求解器的性能,而不是将问题缩小为小内核(例如 SpMM、SpMV)。当使用四种不同的输入矩阵进行测试时,我们的 LOPBCG GPU 实现比优化的 CPU 实现实现了 2.8–4.3 的加速。评估的配置将一台 Skylake CPU 与一台 Skylake CPU 和一台 NVIDIA V100 GPU 进行了比较。我们的 OpenMP 和 OpenACC LOBPCG GPU 实现提供了几乎相同的性能。我们还考虑如何创建一个高效的 LOBPCG 求解器,可以解决大于 GPU 内存容量的问题。为此,我们创建了代表 LOBPCG 中两个主要内核(内积和 SpMM 内核)的微基准,然后评估使用两种不同编程方法时的性能:平铺内核以及将统一内存与原始内核一起使用。与具有 PCIe Gen3 和 NVLink 2.0 CPU 到 GPU 互连的超级计算机上的统一内存实现相比,我们的平铺 SpMM 实现分别实现了 2.9 和 48.2 的加速。
Achieving high performance and performance portability for large-scale scientific applications is a major challenge on heterogeneous computing systems such as many-core CPUs and accelerators like GPUs. In this work, we implement a widely used block eigensolver, Locally Optimal Block Preconditioned Conjugate Gradient (LOBPCG), using two popular directive based programming models (OpenMP and OpenACC) for GPU-accelerated systems. Our work differs from existing work in that it adopts a holistic approach that optimizes the full solver performance rather than narrowing the problem into small kernels (e.g., SpMM, SpMV). Our LOPBCG GPU implementation achieves a 2.8–4.3speedup over an optimized CPU implementation when tested with four different input matrices. The evaluated configuration compared one Skylake CPU to one Skylake CPU and one NVIDIA V100 GPU. Our OpenMP and OpenACC LOBPCG GPU implementations gave nearly identical performance. We also consider how to create an efficient LOBPCG solver that can solve problems larger than GPU memory capacity. To this end, we create microbenchmarks representing the two dominant kernels (inner product and SpMM kernel) in LOBPCG and then evaluate performance when using two different programming approaches: tiling the kernels, and using Unified Memory with the original kernels. Our tiled SpMM implementation achieves a 2.9and 48.2speedup over the Unified Memory implementation on supercomputers with PCIe Gen3 and NVLink 2.0 CPU to GPU interconnects, respectively.
DOI: --
发表时间: 2012
期刊:
影响因子: --
作者:
P. Maris;H. Aktulga;M. Caprio;Ümit V. Çatalyürek;E. Ng;Dossay Oryspayev;H. Potter;Erik Saule;M. Sosonkina;J. Vary;Chao Yang;Zhengguo Zhou
通讯作者: Zhengguo Zhou
DOI: --
发表时间: 2008
期刊: 2008 SC - International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
P. Sternberg;E. Ng;Chao Yang;P. Maris;J. Vary;M. Sosonkina;H. Le
通讯作者: H. Le
多核计算机架构上从头开始核物理计算的扩展
DOI: --
发表时间: 2010
期刊: International Conference on Conceptual Structures
影响因子: --
作者:
P. Maris;M. Sosonkina;J. Vary;E. Ng;Chao Yang
通讯作者: Chao Yang
DOI: 10.1109/ipdps.2014.125
发表时间: 2014-05
期刊: 2014 IEEE 28th International Parallel and Distributed Processing Symposium
影响因子: --
作者:
H. Aktulga;A. Buluç;Samuel Williams;Chao Yang
通讯作者: H. Aktulga;A. Buluç;Samuel Williams;Chao Yang
DOI: 10.1007/978-3-319-96983-1_48
发表时间: 2018-03
期刊: ArXiv
影响因子: --
作者:
Carl Yang;A. Buluç;John Douglas Owens
通讯作者: Carl Yang;A. Buluç;John Douglas Owens