Reinforcement Learning-Based Joint User Scheduling and Link Configuration in Millimeter-Wave Networks

Reinforcement Learning-Based Joint User Scheduling and Link Configuration in Millimeter-Wave Networks
复制标题

DOI:
10.1109/twc.2022.3215922
复制
发表时间:
2022-07
影响因子:
10.4
通讯作者:
Yi Zhang;R. Heath
Yi Zhang;R. Heath
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yi Zhang;R. Heath

文献摘要

被引文献

相似文献

在本文中,我们开发了联合用户调度和三种毫米波链路配置的算法:中继选择,码本优化和毫米波(mmWave)网络中的波束跟踪。我们的目标是设计一个在线控制器,动态调度用户和配置他们的链路,以最小化系统延迟。为了解决这个复杂的调度问题,我们将其建模为一个动态决策过程,并开发了两个基于强化学习的解决方案。第一个解决方案是基于深度强化学习(DRL),它利用近端策略优化来训练基于神经网络的解决方案。针对DRL潜在的高样本复杂度,我们提出了一种基于经验多臂班组(MAB)的解决方案,该方案将决策过程分解为一系列子动作,并利用经典的最大权重调度和汤普森采样来决定这些子动作。我们对提议的解决方案的评估证实了它们在提供可接受的系统延迟方面的有效性。实验还表明,基于drl的方案具有更好的延迟性能,而基于mab的方案具有更快的训练过程。
In this paper, we develop algorithms for joint user scheduling and three types of mmWave link configuration: relay selection, codebook optimization, and beam tracking in millimeter wave (mmWave) networks. Our goal is to design an online controller that dynamically schedules users and configures their links to minimize system delay. To solve this complex scheduling problem, we model it as a dynamic decision-making process and develop two reinforcement learning-based solutions. The first solution is based on deep reinforcement learning (DRL), which leverages the proximal policy optimization to train a neural network-based solution. Due to the potential high sample complexity of DRL, we also propose an empirical multi-armed bandit (MAB)-based solution, which decomposes the decision-making process into a sequential of sub-actions and exploits classic maxweight scheduling and Thompson sampling to decide those sub-actions. Our evaluation of the proposed solutions confirms their effectiveness in providing acceptable system delay. It also shows that the DRL-based solution has better delay performance while the MAB-based solution has a faster training process.