Reliable Reinforcement Learning Based NOMA Schemes for URLLC

Reliable Reinforcement Learning Based NOMA Schemes for URLLC
复制标题

DOI:
10.1109/globecom46510.2021.9685621
复制
发表时间:
2021-12
期刊:
2021 IEEE Global Communications Conference (GLOBECOM)
影响因子:
--
通讯作者:
Waleed Ahsan;Wenqiang Yi;Yuanwei Liu;A. Nallanathan
Waleed Ahsan;Wenqiang Yi;Yuanwei Liu;A. Nallanathan
中科院分区:
其他
文献类型:
--
作者:
Waleed Ahsan;Wenqiang Yi;Yuanwei Liu;A. Nallanathan

文献摘要

相似文献

在本文中,我们提出了一种深度状态-动作-奖励-状态-动作(SARSA)$A$学习方法,用于优化非正交多址接入(NOMA)辅助超可靠低延迟通信(URLLC)中的上行链路资源分配。为了降低时变网络环境中的平均解码错误概率,本文设计了一种可靠的学习算法,用于提供长期的资源分配,其中奖励反馈是基于瞬时网络性能。借助于所提出的算法,本文解决了NOMA-URLLC网络中的可靠资源共享的三个主要挑战:1)动态用户聚类; 2)瞬时反馈系统; 3)最优资源分配。所有这些设计都与所考虑的通信环境相互作用。仿真结果表明:1)与传统Q学习算法相比,该算法收敛速度更快,性能更好; 2)NOMA辅助URLLC系统在译码错误概率方面优于传统OMA系统; 3)动态反馈系统对长期学习过程有效。
In this paper, we propose a deep state-action-reward-state-action (SARSA) $A$ learning approach for optimising the uplink resource allocation in non-orthogonal multiple access (NOMA) aided ultra-reliable low-latency communication (URLLC). To reduce the mean decoding error probability in time-varying network environments, this work designs a reliable learning algorithm for providing a long-term resource allocation, where the reward feedback is based on the instantaneous network performance. With the aid of the proposed algorithm, this paper addresses three main challenges of the reliable resource sharing in NOMA-URLLC networks: 1) Dynamic user clustering; 2) Instantaneous feedback system; and 3) Optimal resource allocation. All of these designs interact with the considered communication environment. The simulation outcomes show that: 1) Compared with the traditional Q learning algorithm, the proposed solution converges faster and obtains better performance; 2) NOMA assisted URLLC outperforms traditional OMA systems in terms of decoding error probabilities; and 3) The dynamic feedback system is efficient for the long-term learning process.