Graphics processing unit-accelerated joint-bitplane belief propagation algorithm in DSC

Graphics processing unit-accelerated joint-bitplane belief propagation algorithm in DSC
复制标题

DSC 中图形处理单元加速的联合位平面置信传播算法

DOI:
10.1007/s11227-016-1736-5
复制
发表时间:
2016-06-01
影响因子:
3.3
通讯作者:
Jeon, Gwanggil
Jeon, Gwanggil
中科院分区:
计算机科学4区
文献类型:
--
作者:
Dai, Yuan;Fang, Yong;Jeon, Gwanggil

文献摘要

被引文献

相似文献

具有非平稳相关性的多进制源可以用单个二进制低密度奇偶校验(LDPC)码进行编码,并在分布式源编码中一起解码。联合位平面置信传播 (JBBP) 是一种适用于多进制源的多个位平面的有用解码算法。然而,它存在计算效率低、执行时间长的缺点。受图形处理单元(GPU)的发展和 JBBP 固有的并行特性的推动,我们提出了一种使用计算统一设备架构编程模型在 GPU 上对 JBBP 算法进行计算密集型处理的新方法。 JBBP的不同节点之间的信念传递采用两种不同的并行模式。发现JBBP的瓶颈在于计算符号节点的整体概率质量函数(pmfs)和比特节点的整体置信度。因此,利用数据分区方法将大型 pmf 数组分割成小块,这些小块可以加载到 L1 缓存而不是全局内存中。选择最佳块大小不仅可以为单个线程分配尽可能大的 L1 缓存,而且可以保证每个流多处理器中存在多个活动扭曲。实验结果表明,当使用长度为 6336(分别为长度为 50,688)的 LDPC 累加(LDPCA)代码来压缩源时,与 CPU 上的原始 C 代码相比,JBBP 解码器在 GPU 上可以实现约 20(分别为 41)的加速。使用更长的 LDPCA 代码将进一步获得更好的性能。此外,并行JBBP还应用于高光谱图像压缩和视频编码中,并表现出良好的加速性能。
TheM-ary source with nonstationary correlation can be encoded with a single binary low-density parity-check (LDPC) code and decoded together in distributed source coding. The joint-bitplane belief propagation (JBBP) is a useful decoding algorithm for multiple bitplanes of anM-ary source. However, it suffers from the drawbacks of low computational efficiency and long execution time. Motivated by the evolution of the Graphics Processing Unit (GPU) and the inherent parallel characteristic of the JBBP, we propose a novel approach for the computationally intensive processing of the JBBP algorithm on GPU using the compute unified device architecture programming model. Two different parallel modes are utilized for the belief passing between different nodes of the JBBP. It is found that the bottlenecks of the JBBP lie in computing the overall probability mass functions (pmfs) of symbol nodes and the overall beliefs of bit nodes. Thus, a data partitioning method is leveraged to split a large array of pmfs into small pieces which can be loaded into L1 cache instead of global memory. The optimal block size is selected which not only assigns as large L1 cache as possible for individual thread, but also guarantees multiple active warps in each stream multiprocessor. Experimental results show that when the length-6336 (length-50,688, resp.) LDPC accumulate (LDPCA) code is used to compress the source, the JBBP decoder can achieve about 20(41, resp.) speedup on GPU compared with the original C code on CPU. Better performance would be further obtained with longer LDPCA codes. Moreover, the parallel JBBP is also applied in hyperspectral image compression and video coding and it shows good speedup performance.