Portable mapping of data parallel programs to OpenCL for heterogeneous systems

Portable mapping of data parallel programs to OpenCL for heterogeneous systems
复制标题

DOI:
10.1109/cgo.2013.6494993
复制
发表时间:
2013-02
期刊:
Proceedings of the 2013 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
影响因子:
--
通讯作者:
Dominik Grewe;Zheng Wang;M. O’Boyle
Dominik Grewe;Zheng Wang;M. O’Boyle
中科院分区:
其他
文献类型:
--
作者:
Dominik Grewe;Zheng Wang;M. O’Boyle

文献摘要

被引文献

相似文献

基于GPU的系统非常有吸引力,因为它们具有很少的成本,因此可以实现这种潜力。 GPU。异类多核。我们方案的关键特征是,它利用现有的转换,尤其多核主机。我们在整个NAS平行基准套件上应用了我们的方法GEFORCE GTX 580和CORE 17/AMD RADEON 7970。我们在顺序基线上平均达到了4.51×和4.20×(143×和67×)的平均速度。比独立专家程序员开发的手工编码的,特定于GPU的OPENCL实现更快。
General purpose GPU based systems are highly attractive as they give potentially massive performance at little cost. Realizing such potential is challenging due to the complexity of programming. This paper presents a compiler based approach to automatically generate optimized OpenCL code from data-parallel OpenMP programs for GPUs. Such an approach brings together the benefits of a clear high level-language (OpenMP) and an emerging standard (OpenCL) for heterogeneous multi-cores. A key feature of our scheme is that it leverages existing transformations, especially data transformations, to improve performance on GPU architectures and uses predictive modeling to automatically determine if it is worthwhile running the OpenCL code on the GPU or OpenMP code on the multi-core host. We applied our approach to the entire NAS parallel benchmark suite and evaluated it on two distinct GPU based systems: Core i7/NVIDIA GeForce GTX 580 and Core 17/AMD Radeon 7970. We achieved average (up to) speedups of 4.51× and 4.20× (143× and 67×) respectively over a sequential baseline. This is, on average, a factor 1.63 and 1.56 times faster than a hand-coded, GPU-specific OpenCL implementation developed by independent expert programmers.