A compression-based memory-efficient optimization for out-of-core GPU stencil computation

A compression-based memory-efficient optimization for out-of-core GPU stencil computation
复制标题

DOI:
10.1007/s11227-023-05103-8
复制
发表时间:
2023-02
期刊:
The Journal of Supercomputing
影响因子:
--
通讯作者:
Jingcheng Shen;Linbo Long;Xin Deng;M. Okita;Fumihiko Ino
Jingcheng Shen;Linbo Long;Xin Deng;M. Okita;Fumihiko Ino
中科院分区:
其他
文献类型:
--
作者:
Jingcheng Shen;Linbo Long;Xin Deng;M. Okita;Fumihiko Ino

文献摘要

相似文献

用于核外模板计算的代码管理超过GPU内存容量的数据。然而,这样的代码需要在CPU和GPU之间频繁传输数据,这通常会影响整体性能。在这项工作中,我们提出了一种基于压缩,内存效率的方法来加速核心外模板代码。首先,将动态压缩技术集成到核外计算中,以减少CPU-GPU数据传输。其次,采用单工作缓冲区策略来减少GPU内存使用,使更多的数据可以存储在GPU上以供重用,从而增加了时间分块步骤。实验结果表明,所提出的方法显着减少了GPU的内存使用量的21%,从而创造了空间的时间块的步骤相比,没有压缩的代码的数量增加了一倍。我们提出的方法已被证明可以帮助高阶,数据传输绑定模板代码实现加速高达单精度浮点格式和高达双精度浮点格式的NVIDIA Tesla V100 GPU上相比,没有压缩的代码。
A code for out-of-core stencil computation manages data that exceeds the memory capacity of a GPU. However, such a code necessitates frequent data transfers between the CPU and GPU, which often impede overall performance. In this work, we propose a compression-based, memory-efficient method to accelerate out-of-core stencil codes. First, an on-the-fly compression technique is integrated into the out-of-core computation to reduce CPU-GPU data transfers. Secondly, a single-working-buffer strategy is employed to reduce the GPU memory usage, enabling more data to be stored on the GPU for reuse, resulting in increased temporal blocking steps. Experimental results demonstrate that the proposed method significantly reduces the GPU memory usage by 21%, thereby creating space for doubling the number of temporal blocking steps compared to the codes without compression. Our proposed method has shown to help the high-order, data-transfer-bound stencil codes achieve speedups up tofor single-precision floating-point format and up tofor double-precision floating-point format on an NVIDIA Tesla V100 GPU in comparison with the codes without compression.