Multi-Agent Deep Reinforcement Learning for Massive Access in 5G and Beyond Ultra-Dense NOMA System

Multi-Agent Deep Reinforcement Learning for Massive Access in 5G and Beyond Ultra-Dense NOMA System
复制标题

DOI:
10.1109/twc.2021.3117859
复制
发表时间:
2021
影响因子:
10.4
通讯作者:
Zhenjiang Shi;Jiajia Liu;Shangwei Zhang;N. Kato
Zhenjiang Shi;Jiajia Liu;Shangwei Zhang;N. Kato
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhenjiang Shi;Jiajia Liu;Shangwei Zhang;N. Kato

文献摘要

相似文献

随着机器类型通信(MTC)的快速发展,未来的通信架构需要同时为人型通信(HTC)和MTC提供独具特色的服务。来自MTC的巨大连接给现有的无线网络带来了严峻的挑战。超高密度网络(UDN)通过密集部署小型基站(SBSS)可以支持海量设备接入,是一种很有前途的候选技术。与传统的单基站无线网络中的资源管理不同,统一数字网络中BS级的资源分配问题更加突出,设备的多样性将使这一问题更加复杂。有鉴于此,我们研究了HTC和MTC共存的UDN中海量接入和资源管理的联合优化问题。考虑到计算的复杂性和可扩展性,我们提出了一种基于多智能体深度强化学习的SBS状态选择方案,每个SBS作为一个智能体,通过与环境的持续交互在活动和空闲之间选择最优状态。此外,采用电力域非正交多址进一步提高系统吞吐量,对HTC和MTC分别采用基于授权和免授权的接入方式,以满足其独特的特点。大量的数值结果从多个角度验证了所提方案的优越性能。
With the rapid development of machine-type communications (MTC), the future communication architecture needs to provide services for both human-type communications (HTC) and MTC with unique characteristics. The huge connections from MTC bring serious challenges to the existing wireless network. Ultra-dense network (UDN), a promising candidate technology, can support massive device access through dense deployment of small base stations (SBSs). Different from the resource management in traditional wireless network with single base station (BS), the resource allocation problem at BS level is more prominent in UDN, and the diversity of devices will make this problem more complicated. In view of this, we investigate the joint optimization of massive access and resource management in the UDN where HTC and MTC coexist. Considering the computational complexity and scalability, we propose a multi-agent deep reinforcement learning based SBS state selection scheme, in which each SBS acts as an agent and selects the optimal state between active and idle by continuously interacting with the environment. In addition, we adopt the power-domain non-orthogonal multiple access to further improve system throughput, and use grant-based and grant-free access manners for HTC and MTC respectively, so as to meet their unique characteristics. Extensive numerical results demonstrate the superior performances of proposed scheme in multiple perspectives.