High throughput low latency LDPC decoding on GPU for SDR systems

High throughput low latency LDPC decoding on GPU for SDR systems
复制标题

DOI:
10.1109/globalsip.2013.6737137
复制
发表时间:
2013-12
期刊:
2013 IEEE Global Conference on Signal and Information Processing
影响因子:
--
通讯作者:
Guohui Wang;Michael Wu;Bei Yin;Joseph R. Cavallaro
Guohui Wang;Michael Wu;Bei Yin;Joseph R. Cavallaro
中科院分区:
其他
文献类型:
--
作者:
Guohui Wang;Michael Wu;Bei Yin;Joseph R. Cavallaro

文献摘要

被引文献

相似文献

在本文中,我们提出了一个高吞吐量和低延迟的LDPC(低密度奇偶校验)解码器的GPUs(图形处理单元)上的实现。现有的基于GPU的LDPC解码器实现存在吞吐量低、延迟长的问题,这使得其无法应用于实际的软件无线电系统中。为了克服这个问题,我们提出了一个并行LDPC解码器的优化技术,包括算法优化,完全合并的内存访问,异步数据传输和多流并发内核执行现代GPU架构。实验结果表明,提出的LDPC解码器达到316 Mbps(在10次迭代)的峰值吞吐量在一个单一的GPU。对于从62.5 Mbps到304.16 Mbps的不同吞吐量要求,解码延迟从0.207 ms到1.266 ms变化,比现有技术的解码延迟低得多。当同时使用四个GPU时,我们实现了1.25 Gbps的聚合峰值吞吐量(10次迭代)。
In this paper, we present a high throughput and low latency LDPC (low-density parity-check) decoder implementation on GPUs (graphics processing units). The existing GPU-based LDPC decoder implementations suffer from low throughput and long latency, which prevent them from being used in practical SDR (software-defined radio) systems. To overcome this problem, we present optimization techniques for a parallel LDPC decoder including algorithm optimization, fully coalesced memory access, asynchronous data transfer and multi-stream concurrent kernel execution for modern GPU architectures. Experimental results demonstrate that the proposed LDPC decoder achieves 316 Mbps (at 10 iterations) peak throughput on a single GPU. The decoding latency, which is much lower than that of the state of the art, varies from 0.207 ms to 1.266 ms for different throughput requirements from 62.5 Mbps to 304.16 Mbps. When using four GPUs concurrently, we achieve an aggregate peak throughput of 1.25 Gbps (at 10 iterations).