Multi-Agent Reinforcement Learning With Measured Difference Reward for Multi-Association in Ultra-Dense mmWave Network

Multi-Agent Reinforcement Learning With Measured Difference Reward for Multi-Association in Ultra-Dense mmWave Network
复制标题

DOI:
10.1109/access.2022.3221455
复制
发表时间:
2022
期刊:
影响因子:
3.9
通讯作者:
Xuebin Li;T. Guo;A. B. Mackenzie
Xuebin Li;T. Guo;A. B. Mackenzie
中科院分区:
计算机科学3区
文献类型:
--
作者:
Xuebin Li;T. Guo;A. B. Mackenzie

文献摘要

相似文献

毫米波通信技术有望在满足无线通信中日益增长的对稀缺带宽的需求方面发挥重要作用。然而,毫米波网络非常容易受到阻塞的影响。因此,需要考虑一些缓解技术,例如多连接。密集部署毫米波基站(MBS)以形成超高密度网络(UDN)也有帮助。由于混合了不同的技术,优化分配资源变得具有挑战性。本文研究了具有较多用户设备(UE)的两层异质超密集网络(HetUDN)中的毫米波用户多关联问题。利用多智能体强化学习对通信环境的适应性,提出了一种解决复杂优化问题的多智能体强化学习框架。该方案考虑了基于毫米波波束分割的多连通性,并利用宏基站进行代理之间的间接协作。特别是,我们借用了一种称为差值奖励(DR)的信用分配技术来处理具有大动作空间的相对较大的MAIL系统,据我们所知,这是第一次将具有差值奖励的MAIL应用于用户关联。此外,所提出的方案是可扩展的,主要是由于固定的观测维度和由UE独立采取的个体动作,从而确保操作独立于MBS和UE的数量。数值结果表明,两种测量DR的MARL方案可以在能量效率和服务质量中断之间取得良好的平衡,而使用扩展DR(EDR)的MARL方案提供了额外的性能改进。
Millimeter Wave (mmWave) communication technology is anticipated to play a vital role in meeting the growing demand for the scarce bandwidth in wireless communications. However, mmWave networks are highly susceptible to blockage. Thus, some mitigation techniques, such as multi-connectivity, need to be considered. Densely deploying mmWave base stations (mBSs) to form an ultra-dense network (UDN) also helps. With a mix of different technologies, optimally allocating resources becomes challenging. In this paper, we study mmWave user multi-association in a two-tier heterogeneous ultra-dense network (HetUDN) with a relatively large number of user equipments (UEs). We propose a framework of multi-agent reinforcement learning (MARL) to tackle the complicated optimization problem, leveraging its adaptivity to the communication environment. The proposed scheme considers mmWave beam-division based multi-connectivity and takes advantage of a macro base station (MBS) for indirect cooperation among agents (UEs). In particular, we borrow a credit-assignment technique called difference reward (DR) to deal with a relatively large MARL system with a large action space, which, to the best of our knowledge, is the first time to apply MARL with DR in user association. Furthermore, the proposed schemes are scalable mainly due to fixed observation dimensions and individual actions taken by UEs independently, ensuring that the operation is independent of the numbers of mBSs and UEs. Numerical results suggest that the two MARL schemes with measured DR could achieve a good balance between energy efficiency and QoS outage, and the one using extended DR (EDR) offers additional performance improvement.