Non-episodic and Heterogeneous Environment in Distributed Multi-agent Reinforcement Learning
Non-episodic and Heterogeneous Environment in Distributed Multi-agent Reinforcement Learning
复制标题
DOI:
10.1109/globecom48099.2022.10000672
复制
发表时间:
2022-12
期刊:
影响因子:
--
通讯作者:
Fenghe Hu;Yansha Deng;H. Aghvami
中科院分区:
文献类型:
--
作者:
Fenghe Hu;Yansha Deng;H. Aghvami
Reinforcement learning (RL) is a efficient intelligent algorithm when solving radio resource management problems in the wireless communication network. However, for large-scale networks with limited centralization (i.e., high latency connection to center-server or capacity-limited backbone), it is not realistic to employ a centralized RL algorithm to perform joint real-time decision-making for the entire network, which calls for scalable algorithm designs. Multi-agent RL, which allows separate local execution of policy, has been applied to large-scale wireless communication areas. However, it has performance issue which largely varies with different system settings. In this paper, we study a multi-agent algorithm for a coordinate multipoint (CoMP) scenario, which requires cooperation between base stations. We show that the common settings of user distribution, the design of reward, and episodic in the environment can significantly ease the learning of the algorithm and obtain beautiful converge results. However, these settings are not realistic in wireless communication. By validating the performance difference between these settings with our algorithm in a coordinate multipoint (CoMP) scenario, we introduce several possible solutions and highlight the necessity of further study in this area.