Low latency group‐sorted QR decomposition algorithm for larger‐scale MIMO systems

Low latency group‐sorted QR decomposition algorithm for larger‐scale MIMO systems
复制标题

适用于大规模 MIMO 系统的低延迟组排序 QR 分解算法

DOI:
10.1049/cmu2.12168
复制
发表时间:
2021
期刊:
影响因子:
1.6
通讯作者:
Yongzhong Li
Yongzhong Li
中科院分区:
计算机科学4区
文献类型:
--
作者:
Lirui Chen;Yu Wang;Zuocheng Xing;Shikai Qiu;Qinglin Wang;Yongzhong Li

文献摘要

相似文献

排序 QR 分解 (SQRD) 已广泛应用于各种多输入多输出 (MIMO) 检测器,其中当涉及大规模 MIMO 情况时,排序过程会产生严重的延迟。本文提出了一种群 SQRD (GSQRD) 算法来缓解大规模 MIMO 系统中通用 SQRD 架构的延迟问题。通过在一个阶段对一组 4 列进行预测排序,GSQRD 可以将分解 1616 个复值矩阵的处理延迟消除 41%。此外,对于分解 128128 个矩阵,这个百分比甚至上升到 68%。为了分析副作用,将 GSQRD 应用于仿真链路中的各种 MIMO 检测器,其对 MIMO 检测的性能下降可以忽略不计。此外,GSQRD是一种硬件友好的算法,因为GSQRD中的除法和平方根运算被转换为乘法以简化硬件实现。基于该算法,还采用65 nm CMOS技术实现了两个相应的硬件架构,其中一个排序组中分别包含2列和4列。这些架构可以在 513 MHz 下工作,分解 1616 个复值矩阵。处理延迟分别为 0.32 秒和 0.26 秒,优于最先进的设计。
Sorted QR decomposition (SQRD) has been extensively adopted for various multiple‐input‐multiple‐output (MIMO) detectors, in which the sorting process incurs severe latency when it comes to larger‐scale MIMO situations. This paper proposes a group‐SQRD (GSQRD) algorithm to alleviate the latency problem of general SQRD architectures for larger‐scale MIMO systems. Via predictively sorting a group of 4 columns at one stage, the GSQRD could eliminate the processing latency by 41% for decomposing 1616 complex‐valued matrices. Additionally, this percentage even rises up to 68% for decomposing 128128 matrices. To analyse the side effects, the GSQRD is applied in various MIMO detectors in a simulation link, which exhibits a negligible performance degradation for MIMO detection. Moreover, GSQRD is a hardware‐friendly algorithm because the division and square root operations in GSQRD are converted to multiplications for simplifying the hardware implementation. Based on this algorithm, two corresponding hardware architectures, which contains 2 and 4 columns respectively in a sorting group, are also implemented with 65‐nm CMOS technology. These architectures can work at 513 MHz to decompose 1616 complex‐valued matrices. The processing latencies are respectively 0.32 and 0.26s, superior to the state‐of‐art designs.