Fully Parallelized LZW Decompression for CUDA-Enabled GPUs

Fully Parallelized LZW Decompression for CUDA-Enabled GPUs
复制标题

适用于支持 CUDA 的 GPU 的完全并行化 LZW 解压缩

DOI:
10.1587/transinf.2016pap0011
复制
发表时间:
2016
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Yasuaki Ito
Yasuaki Ito
中科院分区:
--
文献类型:
--
作者:
Shunji Funasaka;K. Nakano;Yasuaki Ito

文献摘要

被引文献

相似文献

本文的主要贡献是为LZW减压提出一种工程平行算法,并在启用CUDA的GPU中实现它,因为顺序LZW减压会通过在压缩文件中读取一个词典,这是一个不容易的为了并行化。然后,我们具有共享内存的标准平行计算模型。 3072像素存储在GEFORCE GTX 980的全局内存中。另一方面,存储在Intel Core i7的主要内存中的同一图像的顺序LZW减压CPU服用了50.1毫秒。我们评估了三种情况的SSD-GPU数据加载时间压缩,大数据,并行算法,GPU,CUDA
The main contribution of this paper is to present a workoptimal parallel algorithm for LZW decompression and to implement it in a CUDA-enabled GPU. Since sequential LZW decompression creates a dictionary table by reading codes in a compressed file one by one, it is not easy to parallelize it. We first present a work-optimal parallel LZW decompression algorithm on the CREW-PRAM (Concurrent-Read Exclusive-Write Parallel Random Access Machine), which is a standard theoretical parallel computing model with a shared memory. We then go on to present an efficient implementation of this parallel algorithm on a GPU. The experimental results show that our GPU implementation performs LZW decompression in 1.15 milliseconds for a gray scale TIFF image with 4096 × 3072 pixels stored in the global memory of GeForce GTX 980. On the other hand, sequential LZW decompression for the same image stored in the main memory of Intel Core i7 CPU takes 50.1 milliseconds. Thus, our parallel LZW decompression on the global memory of the GPU is 43.6 times faster than a sequential LZW decompression on the main memory of the CPU for this image. To show the applicability of our GPU implementation for LZW decompression, we evaluated the SSD-GPU data loading time for three scenarios. The experimental results show that the scenario using our LZW decompression on the GPU is faster than the others. key words: data compression, big data, parallel algorithm, GPU, CUDA