Supporting Data Shuffle Between Threads in OpenMP

Supporting Data Shuffle Between Threads in OpenMP
复制标题

支持 OpenMP 中线程之间的数据洗牌

DOI:
--
复制
发表时间:
2020
期刊:
International Workshop on OpenMP
影响因子:
--
通讯作者:
Yonghong Yan
Yonghong Yan
中科院分区:
--
文献类型:
--
作者:
Anjia Wang;Xinyao Yi;Yonghong Yan

文献摘要

被引文献

相似文献

NVIDIA和AMD GPU都提供了shuffle或permutation指令,以实现不同线程的私有寄存器之间的直接数据移动。由于它不涉及比直接寄存器访问慢的设备上的共享存储器或全局存储器,数据重排提供了优化数据复制以提高计算性能的机会。然而,shuffle是GPU编程的低级原语(NVIDIA和AMD GPU的经线或通道级)。本文提出了在OpenMP中使用shuffle的两种方法:1)使用shuffle指令实现reduction子句的高性能运行时实现; 2)对OpenMP进行shuffle扩展,使用户可以指定何时以及如何在线程之间移动数据。在我们的实验中使用总和减少和2D模板作为例子,洗牌实现总是提供最好的性能,与其他高性能实现相比,加速比高达2.39倍。与2D模板的标准OpenMP卸载代码相比,我们的shuffle实现提供了高达25倍的上级性能。我们还提供了在NVIDIA GPU上使用共享内存的模拟随机播放的研究,以演示如何在没有原生随机播放支持的硬件上支持此扩展。
Both NVIDIA and AMD GPUs provide shuffle or permutation instructions to enable direct data movement between private registers of different threads. Since it doesn’t involve the shared memory or global memory on the device which are slower than direct register access, data shuffling provides opportunities of optimizing data copy to improve computing performance. However, shuffle is low-level primitive(warp- or lane-level for NVIDIA and AMD GPUs) for GPU programming. It requires advanced knowledge and skills to effectively use it. In this paper, we present two approaches of using shuffle in OpenMP, 1) a high performance runtime implementation of reduction clause using shuffle instruction; and 2) proposed shuffle extension to OpenMP to let users specify when and how the data should be moved between threads. Using sum reduction and 2D stencil as examples in our experiment, the shuffle implementation always delivers the best performance with up to 2.39x speedup compared with other high performance implementation. Compared with standard OpenMP offloading code for 2D stencil, our shuffle implementation delivers superior performance for as many as 25x better. We also provide study of simulated shuffle using shared memory on NVIDIA GPUs to demonstrate how to support this extension on hardware that has no native shuffle support.