Benchmarking and Evaluating Unified Memory for OpenMP GPU Offloading

Benchmarking and Evaluating Unified Memory for OpenMP GPU Offloading
复制标题

OpenMP GPU 卸载的统一内存基准测试和评估

DOI:
10.1145/3148173.3148184
复制
发表时间:
2017
期刊:
Proceedings of the Fourth Workshop on the LLVM Compiler Infrastructure in HPC
影响因子:
--
通讯作者:
B. Chapman
B. Chapman
中科院分区:
--
文献类型:
--
作者:
Alok Mishra;Lingda Li;Martin Kong;H. Finkel;B. Chapman

文献摘要

被引文献

相似文献

最新的OpenMP标准提供了自动设备卸载功能,方便GPU编程。尽管如此,仍然存在许多挑战。其中之一就是最近gpu中引入的统一内存特性。当前和未来的高性能计算系统中的gpu已经增强了对统一内存空间的支持。在这样的系统中,CPU和GPU可以透明地访问彼此的内存,即数据的移动由底层系统软硬件自动管理。在这些系统中,内存超过订阅也是可能的。然而,对于这种机制将如何执行,以及程序员应该如何使用它,我们仍然缺乏足够的知识。我们在Rodinia基准测试套件中修改了几个基准测试代码,以研究OpenMP加速器扩展的行为,并使用它们来探索统一内存在OpenMP上下文中的影响。我们还修改了开源的LLVM编译器,以允许OpenMP程序利用统一内存。我们的评估结果表明,虽然统一内存的性能与常规GPU卸载的性能相当,但对于具有大量数据重用的基准测试,当GPU内存过度订阅时,它会遭受显着的开销。基于这些结果,我们为程序员提供了一些指导方针,以实现统一内存的更好性能。
The latest OpenMP standard offers automatic device offloading capabilities which facilitate GPU programming. Despite this, there remain many challenges. One of these is the unified memory feature introduced in recent GPUs. GPUs in current and future HPC systems have enhanced support for unified memory space. In such systems, CPU and GPU can access each other's memory transparently, that is, the data movement is managed automatically by the underlying system software and hardware. Memory over subscription is also possible in these systems. However, there is a significant lack of knowledge about how this mechanism will perform, and how programmers should use it. We have modified several benchmarks codes, in the Rodinia benchmark suite, to study the behavior of OpenMP accelerator extensions and have used them to explore the impact of unified memory in an OpenMP context. We moreover modified the open source LLVM compiler to allow OpenMP programs to exploit unified memory. The results of our evaluation reveal that, while the performance of unified memory is comparable with that of normal GPU offloading for benchmarks with little data reuse, it suffers from significant overhead when GPU memory is over subcribed for benchmarks with large amount of data reuse. Based on these results, we provide several guidelines for programmers to achieve better performance with unified memory.