Multi-Agent Reinforcement Learning-Based Resource Allocation for UAV Networks

Multi-Agent Reinforcement Learning-Based Resource Allocation for UAV Networks
复制标题

DOI:
10.1109/twc.2019.2935201
复制
发表时间:
2020-02-01
影响因子:
10.4
通讯作者:
Nallanathan, Arumugam
Nallanathan, Arumugam
中科院分区:
计算机科学1区
文献类型:
--
作者:
Cui, Jingjing;Liu, Yuanwei;Nallanathan, Arumugam

文献摘要

被引文献

相似文献

无人驾驶飞行器(uav)能够作为空中基站(BSs)提供成本效益和按需无线通信。本文研究了以最大化长期回报为目标的多无人机通信网络的动态资源分配。更具体地说,每架无人机通过自动选择其通信用户、功率电平和子信道与地面用户通信,而无人机之间没有任何信息交换。为了对环境中的动态和不确定性进行建模,我们将长期资源分配问题描述为最大化预期回报的随机博弈,其中每架无人机都成为一个学习代理,每个资源分配解决方案对应于无人机所采取的一个行动。然后,我们开发了一个多智能体强化学习(MARL)框架,每个智能体根据其局部观察使用学习发现其最佳策略。更具体地说,我们提出了一种智能体独立的方法,所有智能体独立地执行决策算法,但基于q学习共享一个共同的结构。最后,仿真结果表明:1)适当的开采和勘探参数能够提高基于MARL的资源分配算法的性能;2)与无人机间信息完全交换的情况相比,本文提出的MARL算法具有可接受的性能。通过这样做,它在性能增益和信息交换开销之间取得了很好的平衡。
Unmanned aerial vehicles (UAVs) are capable of serving as aerial base stations (BSs) for providing both cost-effective and on-demand wireless communications. This article investigates dynamic resource allocation of multiple UAVs enabled communication networks with the goal of maximizing long-term rewards. More particularly, each UAV communicates with a ground user by automatically selecting its communicating user, power level and subchannel without any information exchange among UAVs. To model the dynamics and uncertainty in environments, we formulate the long-term resource allocation problem as a stochastic game for maximizing the expected rewards, where each UAV becomes a learning agent and each resource allocation solution corresponds to an action taken by the UAVs. Afterwards, we develop a multi-agent reinforcement learning (MARL) framework that each agent discovers its best strategy according to its local observations using learning. More specifically, we propose an agent-independent method, for which all agents conduct a decision algorithm independently but share a common structure based on Q-learning. Finally, simulation results reveal that: 1) appropriate parameters for exploitation and exploration are capable of enhancing the performance of the proposed MARL based resource allocation algorithm; 2) the proposed MARL algorithm provides acceptable performance compared to the case with complete information exchanges among UAVs. By doing so, it strikes a good tradeoff between performance gains and information exchange overheads.