Implementing OpenMP’s SIMD Directive in LLVM’s GPU Runtime

Implementing OpenMP’s SIMD Directive in LLVM’s GPU Runtime
复制标题

在 LLVM GPU 运行时中实施 OpenMP SIMD 指令

DOI:
10.1145/3605573.3605640
复制
发表时间:
2023
期刊:
ACM
影响因子:
--
通讯作者:
Chandrasekaran, Sunita
Chandrasekaran, Sunita
中科院分区:
--
文献类型:
--
作者:
Wright, Eric;Doerfert, Johannes;Tian, Shilei;Chapman, Barbara;Chandrasekaran, Sunita

文献摘要

参考文献

相似文献

GPU 支持三个级别的并行性:线程块、块内的扭曲(或波前)以及扭曲内的线程。某些 GPU 编程模型允许使用所有这三个级别,例如使用团队、并行和 simd 指令进行 OpenMP 卸载。然而,LLVM/OpenMP 不支持 simd,仅使用两个级别:线程块和块内的所有线程。对于具有三个显式并行层的代码,这可能会降低性能,并且可能需要重构应用程序。在这项工作中,我们展示了 LLVM 的 OpenMP GPU 运行时中 OpenMP simd 指令的设计和实现,其中包括以 CPU 为中心和以 GPU 为中心的执行模型。我们使用内核和一些代理应用程序评估我们的原型,结果显示性能提高了 1.3 倍到 3.5 倍,具体取决于内核从这种优化中获得的好处。因此,这项工作使现实世界的应用程序能够暴露出三个显式并行层,以更好地利用 GPU 架构的全部优势。
GPUs support three levels of parallelism: thread blocks, warps (or wavefronts) within a block, and threads within a warp. Some GPU programming models allow the use of all three of these levels, such as OpenMP offloading with the teams, parallel, and simd directives. However LLVM/OpenMP does not support simd and only uses two levels, thread blocks and all threads within a block. For codes with three explicit layers of parallelism this can decrease performance and potentially require restructuring of the application. In this work we present our design and implementation of the OpenMP simd directive in LLVM’s OpenMP GPU runtime, which includes both CPU-centric and GPU-centric execution models. We evaluate our prototype using kernels and a few proxy applications showing a performance improvement ranging from 1.3x to 3.5x depending on the benefit the kernels receives from such an optimization. Thus, this work enables real-world applications with three explicit layers of parallelism to expose to better exploit the full benefits of GPU architecture.
使用现代多核 SIMD 架构的矢量结构扩展 OpenMP*
DOI: 10.1007/978-3-642-30961-8_5
发表时间: 2012
期刊: 2009 11th International Conference on Computer Modelling and Simulation
影响因子: --
作者:
Michael Klemm;A. Duran;Xinmin Tian;Hideki Saito;Diego Caballero;X. Martorell
通讯作者: X. Martorell
将三个应用程序移植到 OpenMP 4.5 的早期经验
DOI: --
发表时间: 2016
期刊: International Workshop on OpenMP
影响因子: --
作者:
I. Karlin;T. Scogland;A. Jacob;S. Antão;Gheorghe;C. Bertolli;B. Supinski;E. Draeger;A. Eichenberger;J. Glosli;Holger E. Jones;A. Kunen;David Poliakoff;D. Richards
通讯作者: D. Richards
使用 OpenACC 指令重构 MPS/芝加哥大学辐射 MHD (MURaM) 模型以实现 GPU/CPU 性能可移植性
DOI: 10.1145/3468267.3470576
发表时间: 2021
期刊: Proceedings of the Platform for Advanced Scientific Computing Conference
影响因子: --
作者:
Eric Wright;D. Przybylski;M. Rempel;Cena Miller;S. Suresh;S. Su;R. Loft;S. Chandrasekaran
通讯作者: S. Chandrasekaran
共同设计 OpenMP GPU 运行时和优化以实现接近零开销的执行
DOI: --
发表时间: 2022
期刊: IEEE International Parallel and Distributed Processing Symposium
影响因子: --
作者:
J. Doerfert;Atmn Patel;Joseph Huber;Shilei Tian;J. M. Diaz;Barbara M. Chapman;G. Georgakoudis
通讯作者: G. Georgakoudis
评估拟议的 OpenMP 5.0 功能对性能、可移植性和生产力的影响
DOI: --
发表时间: 2018
期刊: International Workshop on Performance, Portability and Productivity in HPC
影响因子: --
作者:
S. Pennycook;J. Sewall;J. Hammond
通讯作者: J. Hammond