A Deep Q-Network Based-Resource Allocation Scheme for Massive MIMO-NOMA

A Deep Q-Network Based-Resource Allocation Scheme for Massive MIMO-NOMA
复制标题

DOI:
10.1109/lcomm.2021.3055348
复制
发表时间:
2021-05-01
期刊:
IEEE COMMUNICATIONS LETTERS
影响因子:
--
通讯作者:
Zhang, Jia
Zhang, Jia
中科院分区:
其他
文献类型:
--
作者:
Cao, Yanmei;Zhang, Guomei;Zhang, Jia

文献摘要

被引文献

相似文献

针对大规模多输入多输出(MIMO)-非正交多址(NOMA)系统,提出了一种基于深度Q学习网络(DQN)的资源分配(RA)方案。强化学习(RL)框架被开发来构建用于用户聚类、功率分配和波束形成的迭代优化结构。具体地,DQN被设计为基于在功率分配和波束成形之后计算的奖励项来对用户进行分组。目标是最大化奖励项目,即,系统吞吐量。然后,使用反向传播神经网络(BPNN)来实现功率分配。在BP神经网络的训练过程中,将量化幂集中的穷举搜索结果作为输出标签。仿真实验表明,该方案可以获得很高的系统频谱效率,接近基于用户聚类和功率分配的穷举搜索。
In this letter, a deep Q-learning network (DQN) based resource allocation (RA) scheme is proposed for the massive multiple-input multiple-output (MIMO)- nonorthogonal multiple access (NOMA) systems. The reinforcement learning (RL) frame is developed to build an iterative optimization structure for user clustering, power allocation and beamforming. Specifically, a DQN is designed to group the users based on the reward item calculated after power allocation and beamforming. The objective is to maximize the reward item, i.e., the system throughput. Then, a back propagation neural network (BPNN) is used to realize the power allocation. During the training of BPNN, the exhaustive search results in the quantized power set are taken as the output labels. Simulation experiments show that the proposed scheme can achieve high system spectrum efficiency approximating to the exhaustive search based on user clustering and power allocation.