Deep reinforcement learning empowered joint mode selection and resource allocation for RIS-aided D2D communications

Deep reinforcement learning empowered joint mode selection and resource allocation for RIS-aided D2D communications
复制标题

DOI:
10.1007/s00521-023-08745-0
复制
发表时间:
2023-07
影响因子:
6
通讯作者:
Liang Guo;Jie Jia;Jian Chen;An Du;Xingwei Wang
Liang Guo;Jie Jia;Jian Chen;An Du;Xingwei Wang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Liang Guo;Jie Jia;Jian Chen;An Du;Xingwei Wang

文献摘要

相似文献

设备到设备(D2D)通信因其能够提高系统数据速率和资源利用率而被认为是缓解移动流量爆炸问题的一种有前途的解决方案。研究了一种可重构智能表面(RIS)辅助的移动D2D通信框架,在该框架中部署了RIS以提高通信质量。随着D2D对传输距离的变化,D2D对的模式选择和RIS的相移设计在移动场景中至关重要。因此,我们提出了一个模式选择、信道分配、功率分配和离散相移选择的联合优化问题,以最大化D2D对的平均和数据速率。这个问题还受到用户的最大发射功率和最小数据速率要求的限制,后者是为了保证D2D对的公平性。首先将原序贯决策问题转化为马尔可夫博弈问题来求解这一极具挑战性的优化问题。在此基础上,提出了一种多智能体深度强化学习(MADRL)框架,由多智能体共同决定联合模式选择和资源分配策略。该框架结合了多通路深度Q网络(MP-DQN)算法和衰减型DQN算法来解决优化问题。具体地说,我们对D2D对采用MP-DQN算法来处理离散-连续混合动作空间。此外,RIS代理调用衰减DQN算法来选择离散相移。仿真结果表明,该算法在不同情况下均能收敛。在系统性能方面,基于MADRL的算法优于DQN和深度确定性策略梯度(DDPG)的组合算法。此外,还表明部署RIS可以显著提高D2D对的平均和数据速率,并通过增加反射元件(RE)的数量来进一步提高。
Device-to-device (D2D) communication has been regarded as a promising solution to alleviate the mobile traffic explosion problem for its capabilities of improving system data rate and resource utilization. A reconfigurable intelligent surface (RIS) aided mobile D2D communications framework is investigated, where the RIS is deployed to improve communication quality. As the transmission distance of D2D pairs changes, the mode selection for D2D pairs and the phase shift design for RIS is essential for mobile scenarios. Therefore, we formulate a joint optimization problem of mode selection, channel assignment, power allocation, and discrete phase shift selection to maximize the average sum data rate of D2D pairs. This problem is also constrained by the maximum transmit power and the minimum data rate requirements of users, where the latter is to guarantee the fairness of D2D pairs. We first reformulate the original sequential decision-making problem into a Markov game (MG) problem to solve the challenging optimization. Furthermore, a multi-agent deep reinforcement learning (MADRL) framework is proposed, in which multiple agents cooperatively determine the joint mode selection and resource allocation strategy. The proposed MADRL-based framework combines both the multi-pass deep Q-networks (MP-DQN) algorithm and the decaying DQN algorithm to solve the optimization problem. Specifically, we adopt the MP-DQN algorithm for D2D pairs to handle the hybrid discrete-continuous action space. Moreover, the decaying DQN algorithm is invoked by the RIS agent to select discrete phase shifts. Simulation results demonstrate that the proposed algorithm can converge under different cases. The proposed MADRL-based algorithm outperforms the combination algorithm of DQN and the deep deterministic policy gradient (DDPG) in terms of system performance. Moreover, it is also shown that the average sum data rate of D2D pairs can be significantly improved by deploying the RIS and further enhanced by increasing the number of reflecting elements (REs).