Co-Designing an OpenMP GPU Runtime and Optimizations for Near-Zero Overhead Execution

Co-Designing an OpenMP GPU Runtime and Optimizations for Near-Zero Overhead Execution
复制标题

共同设计 OpenMP GPU 运行时和优化以实现接近零开销的执行

DOI:
--
复制
发表时间:
2022
期刊:
IEEE International Parallel and Distributed Processing Symposium
影响因子:
--
通讯作者:
G. Georgakoudis
G. Georgakoudis
中科院分区:
--
文献类型:
--
作者:
J. Doerfert;Atmn Patel;Joseph Huber;Shilei Tian;J. M. Diaz;Barbara M. Chapman;G. Georgakoudis

文献摘要

参考文献

被引文献

相似文献

GPU加速器在现代HPC系统中无处不在。为了对它们进行编程,用户可以选择特定于供应商的本地编程模型,如CUDA,它提供简单的并行语义,最小的运行时支持,或可移植的替代方案,如OpenMP,它提供丰富的并行语义,并具有广泛的运行时库来支持执行。虽然这种运行时的操作很容易限制性能并消耗资源,但在某种程度上,它被认为是不可避免的开销。在这项工作中,我们提出了一个协同设计的方法来优化应用程序,使用一个专门制作的OpenMP GPU运行时,使大多数用例引起接近零的开销。具体来说,我们的方法公开了运行时语义和状态的编译器,优化有效地消除了抽象和运行时状态的最终二进制文件。在用户提供的假设的帮助下,我们可以进一步优化常见模式,否则会增加资源消耗。我们使用多个HPC代理应用程序和基准测试评估了在LLVM/OpenMP GPU卸载基础设施之上构建的原型。CUDA,原来的OpenMP运行时,我们共同设计的替代品的比较表明,通过我们的方法,性能显着提高,资源消耗显着降低。我们通常可以在不牺牲OpenMP的多功能性和可移植性的情况下紧密匹配CUDA实现。
GPU accelerators are ubiquitous in modern HPC systems. To program them, users have the choice between vendor-specific, native programming models, such as CUDA, which provide simple parallelism semantics with minimal runtime support, or portable alternatives, such as OpenMP, which offer rich parallel semantics and feature an extensive runtime library to support execution. While the operations of such a runtime can easily limit performance and drain resources, it was to some degree regarded an unavoidable overhead. In this work we present a co-design methodology for optimizing applications using a specifically crafted OpenMP GPU runtime such that most use cases induce near-zero overhead. Specifically, our approach exposes runtime semantics and state to the compiler such that optimization effectively eliminating abstractions and runtime state from the final binary. With the help of user provided assumptions we can further optimize common patterns that otherwise increase resource consumption. We evaluated our prototype build on top of the LLVM/OpenMP GPU offloading infrastructure with multiple HPC proxy applications and benchmarks. Comparison of CUDA, the original OpenMP runtime, and our co-designed alternative show that, by our approach, performance is significantly improved and resource consumption is significantly lowered. Oftentimes we can closely match the CUDA implementation without sacrificing the versatility and portability of OpenMP.
跨不同计算机架构的性能可移植性
DOI: 10.1109/p3hpc49587.2019.00006
发表时间: 2019
期刊: --
影响因子: --
作者:
Deakin T
通讯作者: Deakin T