A low-complexity high-speed QR decomposition implementation for MIMO receivers

A low-complexity high-speed QR decomposition implementation for MIMO receivers
复制标题

MIMO 接收器的低复杂度高速 QR 分解实现

DOI:
10.1109/iscas.2009.5117678
复制
发表时间:
2009
期刊:
2009 IEEE International Symposium on Circuits and Systems
影响因子:
--
通讯作者:
P. G. Gulak
P. G. Gulak
中科院分区:
--
文献类型:
--
作者:
Dimpesh Patel;M. Shabany;P. G. Gulak

文献摘要

被引文献

相似文献

QR分解(QRD)是许多MIMO信号检测方案的基本信号处理任务,但是,具有较大尺寸的复杂MIMO通道材料会导致高计算复杂性,因此导致大核心区域或低吞吐量移动通信应用程序涉及快速变化的频道,需要在本文中执行QR分解。 QRD方案使用多维式旋转,家庭转换和常规的二维(2D)Givens旋转的组合来降低总体计算复杂性并实现较高的执行能力。提出了新颖的管道结构,该体系结构使用未滚动管道的绳索处理器迭代以最大化吞吐量和资源利用,同时最大程度地减少了门计数。还提出了4×4 MIMO检测器,在0.13µm CMOS Process指示器中,还提供了主要数据处理模块的架构,即2D,Houseperer 3D和4D/2D。一个4×4复合矩阵和四个更新的4×1复合符号向量,每40个周期,以270 MHz的时钟频率,需要36K门设计达到了最低的处理时间,并且在同一框架上报告了待办事项的最高吞吐量。
QR decomposition (QRD) is an essential signal processing task for many MIMO signal detection schemes. However, decomposition of complex MIMO channel matrices with large dimensions leads to high computational complexity, and hence results in either large core area or low throughput. Moreover, for mobile communication applications that involve fast-varying channels, it is required to perform QR decomposition with low processing latency. In this paper, we propose a hybrid QRD scheme that uses a combination of multi-dimensional Givens rotations, Householder transformations and the conventional two-dimensional (2D) Givens rotations to both reduce the overall computational complexity and achieve higher execution parallelism. To prove the effectiveness of the proposed QRD scheme, a novel pipelined architecture is presented that uses un-rolled pipelined CORDIC processors iteratively to maximize throughput and resource utilization, while minimizing the gate count. The architectures of the main data processing modules, namely the 2D, Householder 3D and 4D/2D configurable pipelined CORDIC processors, are also presented. Synthesis results for a 4×4 MIMO detector in 0.13µm CMOS process indicate that this QRD design computes a 4×4 complex R matrix and four updated 4×1 complex symbol vectors every 40 cycles, at a clock frequency of 270 MHz and requires 36K gates. The proposed design achieves the lowest processing time and the highest throughput reported to-date for the same framework.