A Source-to-Source OpenACC Compiler for CUDA

A Source-to-Source OpenACC Compiler for CUDA
复制标题

用于 CUDA 的源到源 OpenACC 编译器

DOI:
--
复制
发表时间:
2013
期刊:
Euro-Par Workshops
影响因子:
--
通讯作者:
M. Sato
M. Sato
中科院分区:
--
文献类型:
--
作者:
Akihiro Tabuchi;M. Nakao;M. Sato

文献摘要

被引文献

相似文献

OpenACC是一个新的基于指令的编程接口,用于GPGPU等加速器。OpenACC允许程序员将数据和计算卸载到加速器,以简化基于CPU的遗留应用程序的移植过程。在本文中,我们介绍了开源OpenACC编译器的设计和实现,该编译器将带有OpenACC指令的C代码转换为带有CUDA API的C代码,CUDA API是为NVIDIA GPU提供的最广泛使用的GPU编程环境。我们采用源代码到源代码的方法,使用Omni编译器基础设施进行源代码分析和翻译。我们的方法将详细的机器特定代码优化留给成熟的NVIDIA CUDA编译器。实验评估的实施表明,我们的编译器编译的一些并行基准代码实现的速度高达31倍以上的CPU,它是商业实现的竞争力。然而,结果也表明OpenACC程序的优化存在一些问题,例如将迭代分配给GPU线程。
OpenACC is a new directive-based programming interface for accelerators such as GPGPU. OpenACC allows the programmer to express the offloading of data and computations to accelerators to simplify the porting process for legacy CPU-based applications. In this paper, we present the design and implementation of an open-source OpenACC compiler that translates C code with OpenACC directives to C code with the CUDA API, which is the most widely used GPU programming environment provided for NVIDIA GPU. We adopt a source-to-source approach using the Omni compiler infrastructure for source code analysis and translations. Our approach leaves detailed machine-specific code optimization to the mature NVIDIA CUDA compiler. An experimental evaluation of the implementation shows that some parallel benchmark codes compiled by our compiler achieve speeds up to 31 times greater than those of CPUs, and that it is competitive with commercial implementations. However, the results also indicate the optimization of OpenACC programs has several problems, such as assigning iterations to GPU threads.