OpenMP Task Generation for Batched Kernel APIs
OpenMP Task Generation for Batched Kernel APIs
复制标题
批量内核 API 的 OpenMP 任务生成
DOI:
10.1007/978-3-030-28596-8_18
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Sato Mitsuhisa
中科院分区:
文献类型:
--
作者:
Lee Jinpil;Watanabe Yutaka;Sato Mitsuhisa
The demand for calculating many small computation kernels is getting significantly important in the HPC area not only for the traditional numerical applications but also recent machine learning applications. While many-core accelerators such as GPUs are power-efficient compute platforms, a large amount of code modification is required. Batched kernel APIs such as batched BLAS can schedule numerical kernels efficiently on the target hardware while it still needs manual code modification. In this paper, we propose a code translation technique to generate batched kernel APIs in a high-level programming model. We use OpenMP task parallelism to specify dependency among numerical kernels. The user adds thetaskdirectives to specify tasks so that the compiler can recognize numerical kernels. The compiler detects conventional numerical kernels in the code and creates a unique batch ID for each kernel. When the task runtime detects tasks with the same batch ID, they are merged into a batch. The current implementation supports NVIDIA GPUs and batched BLAS in cuBLAS. DGEMM kernels can be detected and translated into batched DGEMM. A trivial DGEMM loop and blocked Cholesky decomposition code are used for performance evaluation. The evaluation result shows that batched DGEMM improves the performance when the matrix size is small and the number of DGEMM kernels is large. The time for DGEMMs in blocked Cholesky decomposition is 4 times faster than sequential execution when using batched DGEMM (matrix, tile size 128), however the overall performance is improved 36% because of task/batch management overhead.