OpenMP Task Generation for Batched Kernel APIs

OpenMP Task Generation for Batched Kernel APIs
复制标题

批量内核 API 的 OpenMP 任务生成

DOI:
10.1007/978-3-030-28596-8_18
复制
发表时间:
2019
期刊:
Lecture Notes in Computer Science book series
影响因子:
--
通讯作者:
Sato Mitsuhisa
Sato Mitsuhisa
中科院分区:
--
文献类型:
--
作者:
Lee Jinpil;Watanabe Yutaka;Sato Mitsuhisa

文献摘要

相似文献

在 HPC 领域,计算许多小型计算内核的需求变得越来越重要,不仅对于传统的数值应用程序,而且对于最近的机器学习应用程序也是如此。虽然 GPU 等多核加速器是节能计算平台,但需要进行大量代码修改。批处理内核 API(例如批处理 BLAS)可以在目标硬件上高效地调度数值内核,同时仍然需要手动修改代码。在本文中,我们提出了一种代码翻译技术,用于在高级编程模型中生成批量内核 API。我们使用 OpenMP 任务并行性来指定数值内核之间的依赖关系。用户添加任务指令来指定任务,以便编译器可以识别数字内核。编译器检测代码中的传统数字内核,并为每个内核创建唯一的批次 ID。当任务运行时检测到具有相同批次 ID 的任务时,会将它们合并为一个批次。当前的实现支持 NVIDIA GPU 和 cuBLAS 中的批处理 BLAS。可以检测 DGEMM 内核并将其转换为批量 DGEMM。使用简单的 DGEMM 循环和阻塞 Cholesky 分解代码进行性能评估。评估结果表明,当矩阵尺寸较小且DGEMM核数量较大时,批量DGEMM可以提高性能。使用批处理 DGEMM(矩阵,图块大小 128)时,分块 Cholesky 分解中的 DGEMM 时间比顺序执行快 4 倍,但由于任务/批处理管理开销,总体性能提高了 36%。
The demand for calculating many small computation kernels is getting significantly important in the HPC area not only for the traditional numerical applications but also recent machine learning applications. While many-core accelerators such as GPUs are power-efficient compute platforms, a large amount of code modification is required. Batched kernel APIs such as batched BLAS can schedule numerical kernels efficiently on the target hardware while it still needs manual code modification. In this paper, we propose a code translation technique to generate batched kernel APIs in a high-level programming model. We use OpenMP task parallelism to specify dependency among numerical kernels. The user adds thetaskdirectives to specify tasks so that the compiler can recognize numerical kernels. The compiler detects conventional numerical kernels in the code and creates a unique batch ID for each kernel. When the task runtime detects tasks with the same batch ID, they are merged into a batch. The current implementation supports NVIDIA GPUs and batched BLAS in cuBLAS. DGEMM kernels can be detected and translated into batched DGEMM. A trivial DGEMM loop and blocked Cholesky decomposition code are used for performance evaluation. The evaluation result shows that batched DGEMM improves the performance when the matrix size is small and the number of DGEMM kernels is large. The time for DGEMMs in blocked Cholesky decomposition is 4 times faster than sequential execution when using batched DGEMM (matrix, tile size 128), however the overall performance is improved 36% because of task/batch management overhead.