Deep Reinforcement Learning for Delay-Sensitive LTE Downlink Scheduling

Deep Reinforcement Learning for Delay-Sensitive LTE Downlink Scheduling
复制标题

DOI:
10.1109/pimrc48278.2020.9217110
复制
发表时间:
2020-08
期刊:
2020 IEEE 31st Annual International Symposium on Personal, Indoor and Mobile Radio Communications
影响因子:
--
通讯作者:
Nikhilesh Sharma;Sen Zhang;Someshwar Rao Somayajula Venkata;Filippo Malandra;Nicholas Mastronarde;Jacob Chakareski
Nikhilesh Sharma;Sen Zhang;Someshwar Rao Somayajula Venkata;Filippo Malandra;Nicholas Mastronarde;Jacob Chakareski
中科院分区:
其他
文献类型:
--
作者:
Nikhilesh Sharma;Sen Zhang;Someshwar Rao Somayajula Venkata;Filippo Malandra;Nicholas Mastronarde;Jacob Chakareski

文献摘要

被引文献

相似文献

我们考虑LTE下行链路调度系统,其中基站向运行延迟敏感应用的用户分配资源块(RB)。我们的目标是找到一个调度策略,最大限度地减少排队延迟所经历的用户。我们制定这个问题作为一个马尔可夫决策过程(MDP),集成了每个RB中的每个用户的信道质量指标(CQI),每个用户的队列状态。为了解决这个涉及高维状态和动作空间的复杂问题,我们提出了一个基于深度强化学习的调度框架,该框架利用深度确定性策略梯度(DDPG)算法来最小化用户经历的排队延迟。我们广泛的实验表明,我们的方法优于国家的最先进的基准方面的平均吞吐量,排队延迟和公平性,实现高达55%的排队延迟比最好的基准。
We consider an LTE downlink scheduling system where a base station allocates resource blocks (RBs) to users running delay-sensitive applications. We aim to find a scheduling policy that minimizes the queuing delay experienced by the users. We formulate this problem as a Markov Decision Process (MDP) that integrates the channel quality indicator (CQI) of each user in each RB, and queue status of each user. To solve this complex problem involving high dimensional state and action spaces, we propose a Deep Reinforcement Learning based scheduling framework that utilizes the Deep Deterministic Policy Gradient (DDPG) algorithm to minimize the queuing delay experienced by the users. Our extensive experiments demonstrate that our approach outperforms state-of-the-art benchmarks in terms of average throughput, queuing delay, and fairness, achieving up to 55% lower queuing delay than the best benchmark.