A Deep Reinforcement Learning-Based Resource Scheduler for Massive MIMO Networks

A Deep Reinforcement Learning-Based Resource Scheduler for Massive MIMO Networks
复制标题

DOI:
10.1109/tmlcn.2023.3313988
复制
发表时间:
2023-03
期刊:
IEEE Transactions on Machine Learning in Communications and Networking
影响因子:
--
通讯作者:
Qing An;Santiago Segarra;C. Dick;A. Sabharwal;Rahman Doost-Mohammady
Qing An;Santiago Segarra;C. Dick;A. Sabharwal;Rahman Doost-Mohammady
中科院分区:
其他
文献类型:
--
作者:
Qing An;Santiago Segarra;C. Dick;A. Sabharwal;Rahman Doost-Mohammady

文献摘要

被引文献

相似文献

在大规模MIMO系统中,大量的天线使得基站可以同时与多个用户通信,并通过多用户波束形成来获得频率资源。然而,高度相关的用户信道可能会极大地阻碍多用户波束成形所能实现的频谱效率。因此,基站在每个时间和频率资源块中调度合适的用户组以在遵守用户之间的公平性约束的同时实现最大频谱效率是至关重要的。本文研究了大规模MIMO系统的资源调度问题,其最优解是NP-难的。受深度强化学习(DRL)用于解决大动作集问题的最新研究成果的启发,基于最新的软行为者-批评者(SAC)DRL模型和K-近邻(KNN)算法,提出了一种面向大规模MIMO的动态调度算法SMART。通过使用真实的海量MIMO信道模型以及来自信道测量实验的真实数据集进行综合仿真,我们证明了所提出的模型在不同的信道条件下的有效性。我们的结果表明,我们提出的模型在频谱效率和公平性方面与最优比例公平(OPT-PF)调度算法非常接近,并且在OPT-PF是计算可行的中等网络规模时,计算复杂度降低了一个数量级以上。我们的结果也表明了我们提出的调度器在具有大量用户和资源块的网络中的可行性和高性能。
The large number of antennas in massive MIMO systems allows the base station to communicate with multiple users at the same time and frequency resource with multi-user beamforming. However, highly correlated user channels could drastically impede the spectral efficiency that multi-user beamforming can achieve. As such, it is critical for the base station to schedule a suitable group of users in each time and frequency resource block to achieve maximum spectral efficiency while adhering to fairness constraints among the users. In this paper, we consider the resource scheduling problem for massive MIMO systems with its optimal solution known to be NP-hard. Inspired by recent achievements in deep reinforcement learning (DRL) to solve problems with large action sets, we propose SMART, a dynamic scheduler for massive MIMO based on the state-of-the-art Soft Actor-Critic (SAC) DRL model and the K-Nearest Neighbors (KNN) algorithm. Through comprehensive simulations using realistic massive MIMO channel models as well as real-world datasets from channel measurement experiments, we demonstrate the effectiveness of our proposed model in various channel conditions. Our results show that our proposed model performs very close to the optimal proportionally fair (Opt-PF) scheduler in terms of spectral efficiency and fairness with more than one order of magnitude lower computational complexity in medium network sizes where Opt-PF is computationally feasible. Our results also show the feasibility and high performance of our proposed scheduler in networks with a large number of users and resource blocks.