Distributed Power Allocation for 6-GHz Unlicensed Spectrum Sharing via Multi-agent Deep Reinforcement Learning

Distributed Power Allocation for 6-GHz Unlicensed Spectrum Sharing via Multi-agent Deep Reinforcement Learning
复制标题

DOI:
10.1109/icit58465.2023.10143125
复制
发表时间:
2023-04
期刊:
2023 IEEE International Conference on Industrial Technology (ICIT)
影响因子:
--
通讯作者:
Xiang Zhang;Arupjyoti Bhuyan;S. Kasera;Mingyue Ji
Xiang Zhang;Arupjyoti Bhuyan;S. Kasera;Mingyue Ji
中科院分区:
其他
文献类型:
--
作者:
Xiang Zhang;Arupjyoti Bhuyan;S. Kasera;Mingyue Ji

文献摘要

相似文献

我们认为,由多个蜂窝运营商的频谱共享的问题。提出了一种基于深度强化学习(DRL)的分布式功率分配方案,该方案采用多智能体深度确定性策略梯度(MA-DDPG)算法。特别地,我们将属于共享相同频带的多个运营商的基站(BS)建模为DRL代理,其以同步的方式同时确定向其调度的用户设备(UE)的发射功率。每个BS的功率决定基于其自身对无线电环境(RF)环境的观察,该RF环境包括从其服务的UE报告的干扰测量以及从其他BS获得的有限量的信息。所提出的方案的一个优点是,它通过将其他BS的动作和观察纳入每个BS自己的评论家,这有助于它获得更准确的整体RF环境的感知,解决了多代理场景中RL的单代理非平稳性问题。一个集中式训练分布式执行的框架被用来训练的政策,其中的批评者是在所有BS的联合行动和观察训练,而每个BS的演员只采取本地观察作为输入,以产生发射功率。在6 GHz U-NII-5频段上的仿真结果表明,该功率分配方案比现有的几种方法具有更好的吞吐量性能。
We consider the problem of spectrum sharing by multiple cellular operators. We propose a novel deep Reinforcement Learning (DRL)-based distributed power allocation scheme which utilizes the multi-agent Deep Deterministic Policy Gradient (MA-DDPG) algorithm. In particular, we model the base stations (BSs) that belong to the multiple operators sharing the same band, as DRL agents that simultaneously determine the transmit powers to their scheduled user equipment (UE) in a synchronized manner. The power decision of each BS is based on its own observation of the radio environment (RF) environment, which consists of interference measurements reported from the UEs it serves, and a limited amount of information obtained from other BSs. One advantage of the proposed scheme is that it addresses the single-agent non-stationarity problem of RL in the multi-agent scenario by incorporating the actions and observations of other BSs into each BS's own critic which helps it to gain a more accurate perception of the overall RF environment. A centralized-training-distributed-execution framework is used to train the policies where the critics are trained over the joint actions and observations of all BSs while the actor of each BS only takes the local observation as input in order to produce the transmit power. Simulation with the 6 GHz Unlicensed National Information Infrastructure (U-NII)-5 band shows that the proposed power allocation scheme can achieve better throughput performance than several state-of-the-art approaches.