A Reliable Reinforcement Learning for Resource Allocation in Uplink NOMA-URLLC Networks

A Reliable Reinforcement Learning for Resource Allocation in Uplink NOMA-URLLC Networks
复制标题

DOI:
10.1109/twc.2022.3144618
复制
发表时间:
2022-01
影响因子:
10.4
通讯作者:
Waleed Ahsan;Wenqiang Yi;Yuanwei Liu;A. Nallanathan
Waleed Ahsan;Wenqiang Yi;Yuanwei Liu;A. Nallanathan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Waleed Ahsan;Wenqiang Yi;Yuanwei Liu;A. Nallanathan

文献摘要

相似文献

在本文中,我们提出了一种深度状态-动作-奖励-状态-动作(SARSA) $\lambda $学习方法来优化非正交多址(NOMA)辅助超可靠低延迟通信(URLLC)中的上行资源分配。为了降低时变网络环境下的平均译码错误概率,本工作设计了一种可靠的学习算法,以提供长期的资源分配,其中奖励反馈基于瞬时网络性能。本文利用该算法解决了NOMA-URLLC网络中资源可靠共享面临的三个主要挑战:1)用户聚类;2)瞬时反馈系统;3)资源优化配置。所有这些设计都与所考虑的通信环境交互。最后,我们将该算法与传统Q-learning和SARSA Q-learning算法的性能进行了比较。仿真结果表明:1)与传统的Q学习算法相比,该算法能够在200集内收敛,长期平均误差低至$10^{-2}$;2)在解码错误概率方面,NOMA辅助URLLC优于传统OMA系统;3)所提出的反馈系统对于长期学习过程是有效的。
In this paper, we propose a deep state-action-reward-state-action (SARSA) $\lambda $ learning approach for optimising the uplink resource allocation in non-orthogonal multiple access (NOMA) aided ultra-reliable low-latency communication (URLLC). To reduce the mean decoding error probability in time-varying network environments, this work designs a reliable learning algorithm for providing a long-term resource allocation, where the reward feedback is based on the instantaneous network performance. With the aid of the proposed algorithm, this paper addresses three main challenges of the reliable resource sharing in NOMA-URLLC networks: 1) user clustering; 2) Instantaneous feedback system; and 3) Optimal resource allocation. All of these designs interact with the considered communication environment. Lastly, we compare the performance of the proposed algorithm with conventional Q-learning and SARSA Q-learning algorithms. The simulation outcomes show that: 1) Compared with the traditional Q learning algorithms, the proposed solution is able to converge within 200 episodes for providing as low as $10^{-2}$ long-term mean error; 2) NOMA assisted URLLC outperforms traditional OMA systems in terms of decoding error probabilities; and 3) The proposed feedback system is efficient for the long-term learning process.