Case Study of Using Kokkos and SYCL as Performance-Portable Frameworks for Milc-Dslash Benchmark on NVIDIA, AMD and Intel GPUs

Case Study of Using Kokkos and SYCL as Performance-Portable Frameworks for Milc-Dslash Benchmark on NVIDIA, AMD and Intel GPUs
复制标题

使用 Kokkos 和 SYCL 作为 NVIDIA、AMD 和 Intel GPU 上 Milc-Dslash 基准测试的性能可移植框架的案例研究

DOI:
--
复制
发表时间:
2021
期刊:
International Workshop on Performance, Portability and Productivity in HPC
影响因子:
--
通讯作者:
C. DeTar
C. DeTar
中科院分区:
--
文献类型:
--
作者:
A. S. Dufek;Rahulkumar Gayatri;Neil Mehta;D. Doerfler;B. Cook;Yasaman Ghadar;C. DeTar

文献摘要

被引文献

相似文献

从2021年6月起,TOP500中排名前十的超级计算机中有六台依靠NVIDIA gpu来实现其峰值计算带宽。随着Aurora、Frontier和El Capitan的发布,英特尔和AMD也进入了为科学计算提供gpu的领域。GPU领域日益多样化的结果是出现了可移植编程模型,如Kokkos、SYCL、OpenCL和OpenMP,这些模型允许应用程序开发人员在不同的硬件架构范围内维护单一源代码。虽然可移植框架试图优化给定体系结构上的计算资源使用,但程序员有责任在可以利用gpu上可用的数千个处理元素的应用程序中公开并行性。在本文中,我们介绍了一种gpu友好的Milc-Dslash并行实现,它在算法中暴露了多个并行层次。Milc-Dslash被设计为具有高度优化的矩阵向量乘法的基准,以测量GPU系统上的资源利用率。Milc-Dslash算法中的并行层次结构使用Kokkos和SYCL编程模型映射到目标硬件上。我们介绍了Milc-Dslash在NVIDIA A100 GPU、AMD MI100 GPU和Intel Gen9 GPU上的Kokkos和SYCL实现所取得的性能。此外,我们将Kokkos和SYCL性能分别与NVIDIA A100 GPU和AMD MI100 GPU上使用CUDA和HIP编程模型编写的版本进行了比较。
Six of the top ten supercomputers in the TOP500 list from June 2021 rely on NVIDIA GPUs to achieve their peak compute bandwidth. With the announcement of Aurora, Frontier, and El Capitan, Intel and AMD have also entered the domain of providing GPUs for scientific computing. A consequence of the increased diversity in the GPU landscape is the emergence of portable programming models such as Kokkos, SYCL, OpenCL, and OpenMP, which allow application developers to maintain a single-source code across a diverse range of hardware architectures. While the portable frameworks try to optimize the compute resource usage on a given architecture, it is the programmers responsibility to expose parallelism in an application that can take advantage of thousands of processing elements available on GPUs. In this paper, we introduce a GPU-friendly parallel implementation of Milc-Dslash that exposes multiple hierarchies of parallelism in the algorithm. Milc-Dslash was designed to serve as a benchmark with highly optimized matrix-vector multiplications to measure the resource utilization on the GPU systems. The parallel hierarchies in the Milc-Dslash algorithm are mapped onto a target hardware using Kokkos and SYCL programming models. We present the performance achieved by Kokkos and SYCL implementations of Milc-Dslash on NVIDIA A100 GPU, AMD MI100 GPU, and Intel Gen9 GPU. Additionally, we compare the Kokkos and SYCL performances with those obtained from the versions written in CUDA and HIP programming models on NVIDIA A100 GPU and AMD MI100 GPU, respectively.