Deep Reinforcement Learning-Based Sum Rate Fairness Trade-Off for Cell-Free mMIMO

Deep Reinforcement Learning-Based Sum Rate Fairness Trade-Off for Cell-Free mMIMO
复制标题

基于深度强化学习的无小区mMIMO和速率公平性权衡

DOI:
10.1109/tvt.2022.3230041
复制
发表时间:
2023-05
影响因子:
6.8
通讯作者:
M. Rahmani;M. Bashar;M. Dehghani;A. Akbari;Pei Xiao;R. Tafazolli;M. Debbah
M. Rahmani;M. Bashar;M. Dehghani;A. Akbari;Pei Xiao;R. Tafazolli;M. Debbah
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. Rahmani;M. Bashar;M. Dehghani;A. Akbari;Pei Xiao;R. Tafazolli;M. Debbah

文献摘要

被引文献

相似文献

研究了具有最大比率组合(MRC)和零形式(ZF)方案的无细胞大量输入多输出的上行链路。考虑了一个权力分配优化问题,其中两个相互冲突的指标,即总和率和公平性,共同优化。由于在大规模降低(LSF)组件方面没有可实现速率的封闭式表达,因此无法使用已知的凸优化方法来解决总和率公平权衡优化问题。为了减轻这个问题,我们提出了两种新方法。对于第一种方法,使用使用和验证方案用于为可实现的速率得出封闭形式的表达式。然后,通过提出的顺序凸近似(SCA)方案对公平优化问题进行迭代解决。对于第二种方法,我们利用LSF系数为双胞胎延迟确定性策略梯度(TD3)的输入,该输入有效地解决了非convex sum rate速率公平性权衡优化问题。接下来,分析提出的方案的复杂性和收敛性。数值结果表明,根据ZF和MRC接收器的总和率和最低用户率,所提出的方法比常规功率控制算法具有优越性。此外,拟议的基于TD3的功率控制的性能比拟议的基于SCA的方法以及分数功率方案更好。
The uplink of a cell-free massive multiple-input multiple-output with maximum-ratio combining (MRC) and zero-forcing (ZF) schemes are investigated. A power allocation optimization problem is considered, where two conflicting metrics, namely the sum rate and fairness, are jointly optimized. As there is no closed-form expression for the achievable rate in terms of the large scale-fading (LSF) components, the sum rate fairness trade-off optimization problem cannot be solved by using known convex optimization methods. To alleviate this problem, we propose two new approaches. For the first approach, a use-and-then-forget scheme is utilized to derive a closed-form expression for the achievable rate. Then, the fairness optimization problem is iteratively solved through the proposed sequential convex approximation (SCA) scheme. For the second approach, we exploit LSF coefficients as inputs of a twin delayed deep deterministic policy gradient (TD3), which efficiently solves the non-convex sum rate fairness trade-off optimization problem. Next, the complexity and convergence properties of the proposed schemes are analyzed. Numerical results demonstrate the superiority of the proposed approaches over conventional power control algorithms in terms of the sum rate and minimum user rate for both the ZF and MRC receivers. Moreover, the proposed TD3-based power control achieves better performance than the proposed SCA-based approach as well as the fractional power scheme.