Transfer Reinforcement Learning for 5G New Radio mmWave Networks

Transfer Reinforcement Learning for 5G New Radio mmWave Networks
复制标题

DOI:
10.1109/twc.2020.3044597
复制
发表时间:
2021-05-01
影响因子:
10.4
通讯作者:
Yanikomeroglu, Halim
Yanikomeroglu, Halim
中科院分区:
计算机科学1区
文献类型:
--
作者:
Elsayed, Medhat;Erol-Kantarci, Melike;Yanikomeroglu, Halim

文献摘要

被引文献

相似文献

在本文中,我们的目标是在5G毫米波(毫米波)通信中,通过采用波束成形和非正交多址(NOMA)技术,以提高网络的聚合速率的干扰缓解。尽管毫米波和NOMA的潜在容量增加,但许多技术挑战可能会阻碍性能的提高。特别地,连续干扰消除(SIC)的性能随着每个波束的用户数量的增加而迅速降低,这导致更高的波束内干扰。此外,相邻小区之间的交叉区域引起波束间小区间干扰。为了减轻这两种干扰水平,除了将用户最佳分配到这些波束之外,波束数量的最佳选择是必不可少的。在本文中,我们解决的问题,联合用户小区的关联和选择的波束数量的目的,最大限度地提高总网络容量。我们提出了三种基于机器学习的算法;转移Q学习(TQL),Q学习和最佳SINR关联与基于密度的空间聚类应用程序与噪声(BSDC)算法,并比较它们在不同场景下的性能。在移动性下,TQL和Q-learning在最高流量负载下比BSDC提高了12%的速率。对于静态场景,Q学习和BSDC优于TQL,但与Q学习相比,TQL实现了约29%的收敛加速。
In this paper, we aim at interference mitigation in 5G millimeter-Wave (mm-Wave) communications by employing beamforming and Non-Orthogonal Multiple Access (NOMA) techniques with the aim of improving network's aggregate rate. Despite the potential capacity gains of mm-Wave and NOMA, many technical challenges might hinder that performance gain. In particular, the performance of Successive Interference Cancellation (SIC) diminishes rapidly as the number of users increases per beam, which leads to higher intra-beam interference. Furthermore, intersection regions between adjacent cells give rise to inter-beam inter-cell interference. To mitigate both interference levels, optimal selection of the number of beams in addition to best allocation of users to those beams is essential. In this paper, we address the problem of joint user-cell association and selection of number of beams for the purpose of maximizing the aggregate network capacity. We propose three machine learning-based algorithms; transfer Q-learning (TQL), Q-learning, and Best SINR association with Density-based Spatial Clustering of Applications with Noise (BSDC) algorithms and compare their performance under different scenarios. Under mobility, TQL and Q-learning demonstrate 12% rate improvement over BSDC at the highest offered traffic load. For stationary scenarios, Q-learning and BSDC outperform TQL, however TQL achieves about 29% convergence speedup compared to Q-learning.