OmpMemOpt: Optimized Memory Movement for Heterogeneous Computing

OmpMemOpt: Optimized Memory Movement for Heterogeneous Computing
复制标题

OmpMemOpt:异构计算的优化内存移动

DOI:
10.1007/978-3-030-57675-2_13
复制
发表时间:
2020
期刊:
European Conference on Parallel Processing (Euro-Par 2020
影响因子:
--
通讯作者:
Sarkar, Vivek
Sarkar, Vivek
中科院分区:
--
文献类型:
--
作者:
Barua, Prithayan;Zhao, Jisheng;Sarkar, Vivek

文献摘要

参考文献

被引文献

相似文献

加速架构和应用的快速发展使得异构计算成为高性能计算的标准。将大量数据移动到加速器的成本是应用程序性能和开发人员生产力方面的一个重要瓶颈。内存管理仍然是一项由专业程序员繁琐地执行的手动任务。在本文中,我们开发了一个编译器分析,以自动化异构计算的内存管理。我们提出了一个优化框架,将检测和删除冗余数据移动的问题转化为部分冗余消除(PRE)问题,并应用惰性代码运动技术来优化这些数据移动。我们选择OpenMP作为底层并行编程模型,并在LLVM工具链中实现了我们的优化框架。我们用十个基准测试对其进行了评估,获得了2.3的几何加速比,并平均减少了主机和GPU之间传输的总字节数的50%。
The fast development of acceleration architectures and applications has made heterogeneous computing the norm for high-performance computing. The cost of high volume data movement to the accelerators is an important bottleneck both in terms of application performance and developer productivity. Memory management is still a manual task performed tediously by expert programmers. In this paper, we develop a compiler analysis to automate memory management for heterogeneous computing. We propose an optimization framework that casts the problem of detection and removal of redundant data movements into a partial redundancy elimination (PRE) problem and applies the lazy code motion technique to optimize these data movements. We chose OpenMP as the underlying parallel programming model and implemented our optimization framework in the LLVM toolchain. We evaluated it with ten benchmarks and obtained a geometric speedup of 2.3, and reduced on average 50% of the total bytes transferred between the host and GPU.
OMPSan:OpenMP 数据映射结构的静态验证
DOI: --
发表时间: 2019
期刊: International Workshop on OpenMP
影响因子: --
作者:
Prithayan Barua;J. Shirako;Whitney Tsang;Jeeva Paudel;Wang Chen;Vivek Sarkar
通讯作者: Vivek Sarkar
负载重用分析:设计和评估
DOI: --
发表时间: 1999
期刊: ACM-SIGPLAN Symposium on Programming Language Design and Implementation
影响因子: --
作者:
Rastislav Bodík;Rajiv Gupta;M. Soffa
通讯作者: M. Soffa
使用消除不必要的数据传输将 OpenMP 设备结构转换为 OpenCL
DOI: --
发表时间: 2016
期刊: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
Junghyun Kim;Yong;Jungho Park;Jaejin Lee
通讯作者: Jaejin Lee
在支持 GPU 的 Apache Spark 上透明地避免冗余数据传输
DOI: --
发表时间: 2018
期刊:
影响因子: --
作者:
Ryo Asai;Masao Okita;Fumihiko Ino;and Kenichi Hagihara
通讯作者: and Kenichi Hagihara
多 GPU 机器的自动数据分配和缓冲区管理
DOI: --
发表时间: 2013
期刊: ACM Transactions on Architecture and Code Optimization (TACO)
影响因子: --
作者:
Thejas Ramashekar;Uday Bondhugula
通讯作者: Uday Bondhugula