Reinforcement Learning for User Clustering in NOMA-Enabled Uplink IoT

Reinforcement Learning for User Clustering in NOMA-Enabled Uplink IoT
复制标题

DOI:
10.1109/iccworkshops49005.2020.9145187
复制
发表时间:
2020-06
期刊:
2020 IEEE International Conference on Communications Workshops (ICC Workshops)
影响因子:
--
通讯作者:
Waleed Ahsan;Wenqiang Yi;Yuanwei Liu;Zhijin Qin;A. Nallanathan
Waleed Ahsan;Wenqiang Yi;Yuanwei Liu;Zhijin Qin;A. Nallanathan
中科院分区:
其他
文献类型:
--
作者:
Waleed Ahsan;Wenqiang Yi;Yuanwei Liu;Zhijin Qin;A. Nallanathan

文献摘要

相似文献

模型驱动的算法在无线通信中已经被研究了几十年。目前,在非正交多址(NOMA)领域,基于机器学习技术的无模型方法正在迅速发展,以动态优化多个参数(例如,资源块的数量和服务质量)。在SARSA Q学习和深度强化学习(DRL)的基础上,提出了一种基于用户聚类的多小区系统上行NOMA资源分配方法。它根据网络流量进行用户分组,以有效地利用可用资源,我们将SARSA Q学习应用于轻流量,将DRL应用于高流量网络。为了表征所提出的优化算法的性能,使用为所有用户实现的容量来定义奖励函数。所提出的SARSA Q学习和DRL算法能够帮助基站在不同的业务条件下有效地为物联网用户分配可用资源。仿真结果表明,无论是SARSA Q学习算法还是DRL算法,在所有的实验中都优于OMA算法,并且都以最大和速率收敛。
The model-driven algorithms have been investigated in wireless communications for decades. Presently, the model-free methods based on machine learning techniques are rapidly being developed in the field of non-orthogonal multiple access (NOMA) to dynamically optimize multiples parameters (e.g., number of resource blocks and QoS). With the aid of SARSA Q-learning and Deep reinforcement Learning (DRL), in this paper, we proposed a user clustering-based resource allocation with uplink NOMA techniques in multi-cell systems. It performs user grouping based on network traffic to efficiently utilise the available resources, we apply SARSA Q-learning to light and DRL to heavy network traffic. To characterize the performance of the proposed optimization algorithms, achieved the capacity for all the users is used to define the reward function. The proposed SARSA Q-learning and DRL algorithms are capable of assisting base-stations to efficiently assign available resources to IoT users considering different traffic conditions. As a result, simulation outcomes show that both the algorithms, SARSA Q-learning and DRL performed better than orthogonal multiple access (OMA) in all the experiments and converged with maximum sum-rate.