Dynamic Spectrum Access in Non-stationary Environments: A DRL-LSTM Integrated Approach

Dynamic Spectrum Access in Non-stationary Environments: A DRL-LSTM Integrated Approach
复制标题

DOI:
10.1109/icnc57223.2023.10074454
复制
发表时间:
2023-02
期刊:
2023 International Conference on Computing, Networking and Communications (ICNC)
影响因子:
--
通讯作者:
Ming Feng;Wenhan Zhang;M. Krunz
Ming Feng;Wenhan Zhang;M. Krunz
中科院分区:
其他
文献类型:
--
作者:
Ming Feng;Wenhan Zhang;M. Krunz

文献摘要

相似文献

本文研究了非平稳环境下的动态频谱接入问题,其中次用户和主用户在一组共享的正交信道上工作。非平稳性是由时变的PU活动和不同SU的耦合信道访问策略引起的。考虑到这种非平稳性和信道的动态性,DSA问题被表示为一个隐模式马尔可夫决策过程(HMMDP),它可以被分解成多个MDP在不同的模式。在每个时间,模式之一是活动的,每个模式对应于一个唯一的MDP。当确定活动模式时求解HMMDP,并且求解该模式下的MDP。我们首先提出了一个深度强化学习(DRL)框架,用于在给定模式下求解MDP。然后,我们提出了一种基于长短期记忆(LSTM)的方法来预测每个时隙的活动模式。仿真结果表明,该方案优于基准方案,实现显着减少冲突和提高频谱利用率。
In this paper, we investigate the problem of dynamic spectrum access (DSA) in non-stationary environments, Where secondary users (SUs) and primary users (PUs) operate over a shared set of orthogonal channels. The non-stationarity is caused by the time-varying PU activity and the coupled channel access strategies of different SUs. Considering such non-stationarity and the channel dynamics, the DSA problem is formulated as a hidden-mode Markov Decision Process (HMMDP), Which can be decomposed into multiple MDPs under different modes. At each time, one of the modes is active, each mode corresponds to a unique MDP. The HMMDP is solved when the active mode is determined and the MDP under this mode is solved. We first propose a deep reinforcement learning (DRL) framework for solving the MDP under a given mode. We then propose a long short-term memory (LSTM)-based approach to predict the active mode at each time slot. Simulation results show that the proposed scheme outperforms benchmark schemes by achieving significantly fewer collisions and improved spectrum utilization.