Tile-based Lightweight Integer Compression in GPU

Tile-based Lightweight Integer Compression in GPU
复制标题

DOI:
10.1145/3514221.3526132
复制
发表时间:
2022-06
期刊:
Proceedings of the 2022 International Conference on Management of Data
影响因子:
--
通讯作者:
Anil Shanbhag;B. Yogatama;Xiangyao Yu;S. Madden
Anil Shanbhag;B. Yogatama;Xiangyao Yu;S. Madden
中科院分区:
其他
文献类型:
--
作者:
Anil Shanbhag;B. Yogatama;Xiangyao Yu;S. Madden

文献摘要

相似文献

由于GPU能够使用大规模并行来加速计算,因此越来越多地将其用于高性能和交互式数据分析工作负载。目前,基于GPU的数据分析的一个关键制约因素是GPU设备的内存容量有限。数据压缩是一种强大的技术,可以通过两种方式缓解容量限制:(1)将更多数据放入GPU内存中;(2)加快CPU和GPU之间的数据传输。然而,目前用于GPU的压缩方案在压缩比和/或解压缩速度方面仍然受到限制。我们确定了现有方法的两个限制因素。首先,现有的解压缩解决方案需要多次扫描全局内存以解码多层压缩方案,这会导致大量的内存流量并损害性能。我们提出了基于块的解压缩模型,可以在全局内存上一次完成编码数据的解压缩,并内联查询执行。其次,我们在GPU环境下开发了一种基于比特打包的压缩方案及其优化技术的高效实现。我们的评估表明,我们的方案可以获得与GPU中最好的压缩方案(即nvCOMP)相似的压缩比,而在解压缩速度和查询运行时间方面分别快2.2倍和2.6倍。
GPUs are increasingly used for high-performance and interactive data analytics workloads due to their capability to accelerate computation using massive parallelism. A key constraint of GPU-based data analytics today is the limited memory capacity in GPU devices. Data compression is a powerful technique that can mitigate the capacity limitation in two ways: (1) fitting more data into GPU memory and (2) speeding up data transfer between CPU and GPU. However, compression schemes for GPU today are still limited in compression ratio and/or decompression speed. We identify two limiting factors of existing approaches. First, existing decompression solutions require multiple passes of scanning the global memory to decode layers of compression schemes, incurring significant memory traffic and hurting performance. We present the tile-based decompression model to decompress encoded data in a single pass over global memory and inline with query execution. Second, we develop an efficient implementation of bit-packing-based compression schemes and their optimization techniques in the context of GPU. Our evaluation shows that our schemes can achieve similar compression rates to the best state-of-the-art compression schemes in GPU (i.e., nvCOMP) while being 2.2× and 2.6× faster in decompression speed and query running time.