Non-episodic and Heterogeneous Environment in Distributed Multi-agent Reinforcement Learning

Non-episodic and Heterogeneous Environment in Distributed Multi-agent Reinforcement Learning
复制标题

DOI:
10.1109/globecom48099.2022.10000672
复制
发表时间:
2022-12
期刊:
GLOBECOM 2022 - 2022 IEEE Global Communications Conference
影响因子:
--
通讯作者:
Fenghe Hu;Yansha Deng;H. Aghvami
Fenghe Hu;Yansha Deng;H. Aghvami
中科院分区:
其他
文献类型:
--
作者:
Fenghe Hu;Yansha Deng;H. Aghvami

文献摘要

相似文献

强化学习是解决无线通信网络中无线资源管理问题的一种有效的智能算法。然而,对于具有有限集中性的大规模网络(即,高延迟连接到中心服务器或容量有限的骨干网),采用集中式RL算法来为整个网络执行联合实时决策是不现实的,这需要可扩展的算法设计。多Agent强化学习允许策略的局部执行,已被应用于大规模无线通信领域。然而,它有性能问题,这在很大程度上取决于不同的系统设置。在本文中,我们研究了一个多代理算法的协调多点(CoMP)的情况下,这需要基站之间的合作。我们表明,用户分布,奖励的设计,并在环境中情节的共同设置可以显着减轻学习的算法,并获得漂亮的收敛结果。然而,这些设置在无线通信中是不现实的。通过验证这些设置与我们的算法在一个坐标多点(CoMP)的情况下的性能差异,我们介绍了几种可能的解决方案,并强调在这方面进一步研究的必要性。
Reinforcement learning (RL) is a efficient intelligent algorithm when solving radio resource management problems in the wireless communication network. However, for large-scale networks with limited centralization (i.e., high latency connection to center-server or capacity-limited backbone), it is not realistic to employ a centralized RL algorithm to perform joint real-time decision-making for the entire network, which calls for scalable algorithm designs. Multi-agent RL, which allows separate local execution of policy, has been applied to large-scale wireless communication areas. However, it has performance issue which largely varies with different system settings. In this paper, we study a multi-agent algorithm for a coordinate multipoint (CoMP) scenario, which requires cooperation between base stations. We show that the common settings of user distribution, the design of reward, and episodic in the environment can significantly ease the learning of the algorithm and obtain beautiful converge results. However, these settings are not realistic in wireless communication. By validating the performance difference between these settings with our algorithm in a coordinate multipoint (CoMP) scenario, we introduce several possible solutions and highlight the necessity of further study in this area.