Lightweight Huffman Coding for Efficient GPU Compression

Lightweight Huffman Coding for Efficient GPU Compression
复制标题

DOI:
10.1145/3577193.3593736
复制
发表时间:
2023-06
期刊:
Proceedings of the 37th International Conference on Supercomputing
影响因子:
--
通讯作者:
Milan Shah;Xiaodong Yu;S. Di;M. Becchi;F. Cappello
Milan Shah;Xiaodong Yu;S. Di;M. Becchi;F. Cappello
中科院分区:
其他
文献类型:
--
作者:
Milan Shah;Xiaodong Yu;S. Di;M. Becchi;F. Cappello

文献摘要

相似文献

经常在科学应用程序中部署有损失的压缩,以减少数据传输,尤其是对于需要在飞行的压缩的应用程序。提高基于GPU的损失压缩机Cusz的性能。尤其是对于较小的数据集,我们的工作旨在减少对Cusz压缩效率的霍夫曼编码时间。输入Cusz的Huffman编码阶段。我们没有计算自定义代码簿。 A100 GPU我们的评估表明,我们的方法可以将霍夫曼编码的惩罚降低为78--92倍,转化为高达基线cusz的总速度HDF5块尺寸在压缩的8×加速度上享受,而MB的MPI消息的尺寸为数十MB的尺寸,通信时间的速度为1.4---30.5倍。
Lossy compression is often deployed in scientific applications to reduce data footprint and improve data transfers and I/O performance. Especially for applications requiring on-the-flight compression, it is essential to minimize compression's runtime. In this paper, we design a scheme to improve the performance of cuSZ, a GPU-based lossy compressor. We observe that Huffman coding - used by cuSZ to compress metadata generated during compression - incurs a performance overhead that can be significant, especially for smaller datasets. Our work seeks to reduce the Huffman coding runtime with minimal-to-no impact on cuSZ's compression efficiency. Our contributions are as follows. First, we examine a variety of probability distributions to determine which distributions closely model the input to cuSZ's Huffman coding stage. From these distributions, we create a dictionary of pre-computed codebooks such that during compression, a codebook is selected from the dictionary instead of computing a custom codebook. Second, we explore three codebook selection criteria to be applied at runtime. Finally, we evaluate our scheme on real-world datasets and in the context of two important application use cases, HDF5 and MPI, using an NVIDIA A100 GPU. Our evaluation shows that our method can reduce the Huffman coding penalty by a factor of 78--92×, translating to a total speedup of up to 5× over baseline cuSZ. Smaller HDF5 chunk sizes enjoy over an 8× speedup in compression and MPI messages on the scale of tens of MB have a 1.4--30.5× speedup in communication time.