An OpenACC extension for data layout transformation.

An OpenACC extension for data layout transformation.
复制标题

用于数据布局转换的 OpenACC 扩展。

DOI:
10.1109/waccpd.2014.12
复制
发表时间:
2014
期刊:
WACCPD '14 Proceedings of the First Workshop on Accelerator Programming using Directives
影响因子:
--
通讯作者:
and Satoshi Matsuoka
and Satoshi Matsuoka
中科院分区:
--
文献类型:
--
作者:
Tetsuya Hoshino;Naoya Maruyama;and Satoshi Matsuoka

文献摘要

相似文献

OpenACC作为一种隐式和可移植的接口,在将传统的基于CPU的应用程序移植到涉及GPU和英特尔至强融核等众核加速器的异构、高度并行的计算环境中,正获得越来越大的发展势头。OpenACC提供了一组类似于OpenMP的循环指令,用于并行化和管理数据移动,实现跨不同异构设备的功能可移植性;然而,OpenACC的性能可移植性据说由于不同目标设备的特性而不足,特别是关于内存布局的那些,因为编译器自动尝试适应目前很困难。我们目前正在努力提出一组指令,以允许编译器有更好的语义信息进行调整;在这里,我们特别关注数据布局,如数组结构,这是GPU的有利数据结构,而不是结构数组,它在CPU上表现出良好的性能。我们提出了一个指令扩展OpenACC,允许用户灵活地指定最佳布局,即使数据结构是嵌套的。性能结果表明,与没有此类指令的程序相比,我们的CPU性能提高了96%,GPU性能提高了165%,基本上实现了OpenACC中的功能和性能可移植性。
OpenACC is gaining momentum as an implicit and portable interface in porting legacy CPU-based applications to heterogeneous, highly parallel computational environment involving many-core accelerators such as GPUs and Intel Xeon Phi. OpenACC provides a set of loop directives similar to OpenMP for the parallelization and also to manage data movement, attaining functional portability across different heterogeneous devices; however, the performance portability of OpenACC is said to be insufficient due to the characteristics of different target devices, especially those regarding memory layouts, as automated attempts by the compilers to adapt is currently difficult. We are currently working to propose a set of directives to allow compilers to have better semantic information for adaptation; here, we particularly focus on data layout such as Structure of Arrays, advantageous data structure for GPUs, as opposed to Array of Structures, which exhibits good performance on CPUs. We propose a directive extension to OpenACC that allows the users to flexibility specify optimal layouts, even if the data structures are nested. Performance results show that we gain as much as 96 % in performance for CPUs and 165% for GPUs compared to programs without such directives, essentially attaining both functional and performance portability in OpenACC.