Reinforcement Learning of Beam Codebooks in Millimeter Wave and Terahertz MIMO Systems

Reinforcement Learning of Beam Codebooks in Millimeter Wave and Terahertz MIMO Systems
复制标题

毫米波和太赫兹MIMO系统中波束码本的强化学习

DOI:
10.1109/tcomm.2021.3126856
复制
发表时间:
2021-02
影响因子:
8.3
通讯作者:
Yu Zhang;Muhammad Alrabeiah;A. Alkhateeb
Yu Zhang;Muhammad Alrabeiah;A. Alkhateeb
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yu Zhang;Muhammad Alrabeiah;A. Alkhateeb

文献摘要

被引文献

相似文献

毫米波(mmWave)和太赫兹MIMO系统依赖于预定义的波束成形码本来进行初始接入和数据传输。然而,这些预定义的码本通常没有针对特定环境、用户分布和/或可能的硬件损伤进行优化。这导致具有高波束训练开销的大码本大小,这使得这些系统难以支持高度移动的应用。为了克服这些限制,本文开发了一个深度强化学习框架,该框架学习如何仅依赖于接收功率测量来优化码本波束模式。开发的模型学习如何根据周围环境,用户分布,硬件损伤和阵列几何形状来调整波束方向图。此外,该方法不需要关于信道、RF硬件或用户位置的任何知识。为了减少学习时间,该模型设计了一种新的Wolpertinger变体架构,能够有效地搜索大的离散动作空间。所提出的学习框架尊重RF硬件约束,如恒模和量化移相器的约束。仿真结果证实了所开发的框架的能力,学习接近最佳的波束方向图的视线(LOS),非LOS(NLOS),混合LOS/NLOS的情况下,并与硬件损伤的阵列,而不需要任何信道知识。
Millimeter wave (mmWave) and terahertz MIMO systems rely on pre-defined beamforming codebooks for both initial access and data transmission. These pre-defined codebooks, however, are commonly not optimized for specific environments, user distributions, and/or possible hardware impairments. This leads to large codebook sizes with high beam training overhead which makes it hard for these systems to support highly mobile applications. To overcome these limitations, this paper develops a deep reinforcement learning framework that learns how to optimize the codebook beam patterns relying only on the receive power measurements. The developed model learns how to adapt the beam patterns based on the surrounding environment, user distribution, hardware impairments, and array geometry. Further, this approach does not require any knowledge about the channel, RF hardware, or user positions. To reduce the learning time, the proposed model designs a novel Wolpertinger-variant architecture that is capable of efficiently searching the large discrete action space. The proposed learning framework respects the RF hardware constraints such as the constant-modulus and quantized phase shifter constraints. Simulation results confirm the ability of the developed framework to learn near-optimal beam patterns for line-of-sight (LOS), non-LOS (NLOS), mixed LOS/NLOS scenarios and for arrays with hardware impairments without requiring any channel knowledge.