Distributed Reinforcement Learning Based Delay Sensitive Decentralized Resource Scheduling

Distributed Reinforcement Learning Based Delay Sensitive Decentralized Resource Scheduling
复制标题

DOI:
10.23919/wiopt58741.2023.10349880
复制
发表时间:
2023-08
期刊:
2023 21st International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks (WiOpt)
影响因子:
--
通讯作者:
Geetha Chandrasekaran;G. Veciana
Geetha Chandrasekaran;G. Veciana
中科院分区:
其他
文献类型:
--
作者:
Geetha Chandrasekaran;G. Veciana

文献摘要

相似文献

我们解决的问题,分布式资源分配在无线系统中的存在下,动态的用户流量和耦合造成的干扰。我们提出了一个强化学习(RL)框架的基础上分离的频率重用干扰缓解和机会用户调度之间的关注。特别是,我们探讨了一个设置在基站之间建立一个随机游戏,学习频率复用模式,并使用多代理RL解决用户调度的基本选择。我们建立的存在性和收敛到纳什均衡的建议设置。我们的框架和理论研究结果的性能进行评估,通过模拟和比较更积极的甲骨文辅助集中基线。所得到的频率重用政策,以实现5-25%的容量和相关的延迟性能的改善,在一个集中式的干扰感知最大权重调度策略跨BS。此外,与集中式基准相比,减少9-34%的物理资源利用率导致更高的能源效率。
We address the problem of distributed resource allocation in wireless systems in the presence of dynamic user traffic and coupling resulting from interference. We propose a Reinforcement Learning (RL) framework based on a separation of concerns between frequency reuse for interference mitigation and opportunistic user scheduling. In particular we explore a setting where a stochastic game is set up among base stations to learn frequency reuse patterns and solved using multi-agent RL given an underlying choice for user scheduling. We establish the existence and convergence to a Nash equilibrium of the proposed setting. The performance of our framework and theoretical findings are evaluated through simulation and compared to more aggressive oracle-aided centralized baselines. The resulting frequency reuse policy is shown to achieve 5–25% improvements in capacity and associated delay performance over a centralized interference aware max weight scheduling policy across BSs. Furthermore, a reduced physical resource utilization on the order of 9–34% leads to a higher energy efficiency as compared to the centralized benchmark.