Deep Reinforcement Learning Optimal Transmission Algorithm for Cognitive Internet of Things With RF Energy Harvesting

Deep Reinforcement Learning Optimal Transmission Algorithm for Cognitive Internet of Things With RF Energy Harvesting
复制标题

DOI:
10.1109/tccn.2022.3142727
复制
发表时间:
2022-06-01
影响因子:
8.6
通讯作者:
Zhao, Xiaohui
Zhao, Xiaohui
中科院分区:
计算机科学2区
文献类型:
--
作者:
Guo, Shaoai;Zhao, Xiaohui

文献摘要

被引文献

相似文献

频谱的稀缺性和能量的有限性已经成为物联网设计中的两个关键问题。认知无线电(CR)和射频(RF)能量收集作为两种很有前途的技术,可以共同使用以提高能量和频谱效率。本文研究了具有射频能量收集能力的认知物联网(CIoT)中的最优传输问题,其中优化问题被表述为不含任何优先级知识的马尔可夫决策过程(MDP)。针对主用户网络(PUN)的信道活动状态、射频能量到达过程和信道信息无法提前获取的情况,提出了一种基于深度强化学习(DRL)的深度确定性策略梯度(DDPG)算法来处理动态上行接入、工作模式选择和连续功率分配,以实现长期上行吞吐量最大化。仿真结果表明,与基于深度q -网络(deep Q-network, DQN)的算法、近视算法和随机算法相比,该算法是有效且高效的。
Spectrum scarcity and energy limitation are becoming two critical issues in designing Internet of Things (IoT). As two promising technologies, cognitive radio (CR) and radio frequency (RF) energy harvesting can be used together to improve both energy and spectral efficiency. In this paper, an optimal transmission problem in a cognitive IoT (CIoT) with RF energy harvesting capability is investigated, where the optimization problem is formulated as a Markov decision process (MDP) without any priori-knowledge. Considering that the channel activity states of primary user network (PUN), RF energy arrival process and channel information are not available in advance, a deep reinforcement learning (DRL) based deep deterministic policy gradient (DDPG) algorithm is proposed to deal with the dynamic uplink access, working mode selection and continuous power allocation to maximize a long term uplink throughput. The simulation results show that the proposed algorithm is valid and efficient to achieve better performances when compared with deep Q-network (DQN) based, myopic and random algorithms.