Multi-Agent Reinforcement Learning for Wireless User Scheduling: Performance, Scalablility, and Generalization

Multi-Agent Reinforcement Learning for Wireless User Scheduling: Performance, Scalablility, and Generalization
复制标题

DOI:
10.1109/ieeeconf56349.2022.10051992
复制
发表时间:
2022-10
期刊:
2022 56th Asilomar Conference on Signals, Systems, and Computers
影响因子:
--
通讯作者:
Kun Yang;Donghao Li;Cong Shen;Jing Yang;Shu-ping Yeh;J. Sydir
Kun Yang;Donghao Li;Cong Shen;Jing Yang;Shu-ping Yeh;J. Sydir
中科院分区:
其他
文献类型:
--
作者:
Kun Yang;Donghao Li;Cong Shen;Jing Yang;Shu-ping Yeh;J. Sydir

文献摘要

相似文献

针对蜂窝网络中的用户调度问题,提出了一种多智能体强化学习(MARL)解决方案。结合这个特殊用例的特征,我们将问题置于一个分散的部分可观察马尔可夫决策过程(Dec-POMDP)框架中,并提出了一个允许完全分散执行的马尔可夫决策过程的详细设计。在系统级仿真中全面评估了MARL对集中式强化学习和工程启发式解决方案的性能。特别是,MARL实现了与集中式RL几乎相同的总系统奖励,同时随着基站数量的增加而具有更好的可伸缩性。MARL和集中式强化学习在新环境中的可移植性也进行了研究,并且基于在环境池中训练的通用模型的简单微调方法被证明具有更快的收敛速度,同时与单独训练的强化学习代理实现相当的性能,证明了其泛化能力。
We propose a multi-agent reinforcement learning (MARL) solution for the user scheduling problem in cellular networks. Incorporating features of this particular use case, we cast the problem in a decentralized partially observable Markov decision process (Dec-POMDP) framework, and present a detailed design of MARL that allows for fully decentralized execution. The performance of MARL against both centralized RL and an engineering heuristic solution is comprehensively evaluated in a system-level simulation. In particular, MARL achieves almost the same total system reward as centralized RL, while enjoying much better scalability with the number of base stations. The transferability of both MARL and centralized RL to new environments is also investigated, and a simple fine-tuning approach based on a general model trained on a pool of environments is shown to have faster convergence while achieving comparable performance with individually trained RL agents, demonstrating its generalization capability.