Truly Distributed Multicell Multi-Band Multiuser MIMO by Synergizing Game Theory and Deep Learning

Truly Distributed Multicell Multi-Band Multiuser MIMO by Synergizing Game Theory and Deep Learning
复制标题

DOI:
10.1109/access.2021.3059587
复制
发表时间:
2021
期刊:
影响因子:
3.9
通讯作者:
Kai‐Kit Wong;Guochen Liu;Wenjing Cun;Wenkai Zhang;Mingming Zhao;Zhongbin Zheng
Kai‐Kit Wong;Guochen Liu;Wenjing Cun;Wenkai Zhang;Mingming Zhao;Zhongbin Zheng
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kai‐Kit Wong;Guochen Liu;Wenjing Cun;Wenkai Zhang;Mingming Zhao;Zhongbin Zheng

文献摘要

相似文献

具有大规模多输入多输出(MIMO)的动态频率分配(DFA)是用于多小区通信的有希望的候选方案,其中采用大规模MIMO来最大化每个小区的容量,而通过DFA来解决小区间干扰(ICI)。然而,由于在小区级中的基站(BS)处缺乏可用的全局信道状态,以分布式方式实现该方法是非常困难的。我们利用一个前瞻性的游戏,以自动协调DFA在小区之间的分布式方式,而迫零(ZF)在每个单元格使用,以最大限度地提高复用增益。为了最大限度地提高网络容量,使用离线集中式训练的多智能体深度强化学习(DRL)被用来训练BS掌握其博弈论和解策略。其结果是每个BS都有一个经过训练的神经网络,使其具有与其他BS协调的丰富经验,以收敛到网络有效的均衡。在线算法是分布式的,BS作为专家玩家竞争,使用他们的训练动作开始谈判过程。仿真结果表明,所提出的协同深度学习博弈论算法明显优于仅DRL和仅博弈论方法以及其他多小区MIMO基准。
Dynamic frequency allocation (DFA) with massive multiple-input multiple-output (MIMO) is a promising candidate for multicell communications where massive MIMO is adopted to maximize the per-cell capacity whereas the inter-cell interference (ICI) is tackled by DFA. Realizing this approach in a distributed fashion is however very difficult due to the lack of global channel state available at the base stations (BSs) in the cell level. We utilize a forward-looking game to automate reconciliation for DFA in a distributed manner between cells while zero-forcing (ZF) is used at each cell to maximize the multiplexing gain. To maximize the network capacity, multi-agent deep reinforcement learning (DRL) using offline centralized training is leveraged to train the BSs to master their game-theoretic reconciliation strategies. The result is a trained neural network for each BS, empowering it with rich experience of reconciliation with other BSs for converging to a network-efficient equilibrium. The online algorithm is distributed with the BSs competing as expert players to start the negotiation process using their trained actions. Simulation results show that the proposed synergized deep-learning game-theoretic algorithm outperforms significantly the DRL-only and game-theoretic only methods, and other benchmarks for multicell MIMO.