Adaptive UAV-Trajectory Optimization Under Quality of Service Constraints: A Model-Free Solution

Adaptive UAV-Trajectory Optimization Under Quality of Service Constraints: A Model-Free Solution
复制标题

DOI:
10.1109/access.2020.3001752
复制
发表时间:
2020-01-01
期刊:
影响因子:
3.9
通讯作者:
Hanzo, Lajos
Hanzo, Lajos
中科院分区:
计算机科学3区
文献类型:
--
作者:
Cui, Jingjing;Ding, Zhiguo;Hanzo, Lajos

文献摘要

被引文献

相似文献

无人机(UAV)具有提供可靠的高速连接的潜力,正在成为未来无线网络的一个有前途的组成部分。无人机从一组随机分布的传感器收集数据,这些传感器的位置和要传输的数据量对无人机来说都是未知的。为了帮助无人机在没有上述知识的情况下在不确定的情况下找到最优的运动轨迹,同时以最大化累积收集的数据为目标,以无人机为学习主体,将运动轨迹建模为马尔可夫决策过程,建立了一个强化学习问题。在此基础上,提出了两种新的基于随机建模和强化学习的航迹优化算法,使无人机在不需要系统辨识的情况下进行航迹优化。更具体地说,通过将考虑的区域划分为小块,我们提出了基于状态-动作-奖励-状态动作(SASA)和Q-学习的无人机轨迹优化算法(即SUTOA和QUTOA),旨在最大化在有限飞行时间内收集的累积数据。仿真结果表明,两种方法都能在飞行时间约束下找到最优航迹。QUTOA和SUTOA的偏好取决于无人机起点和终点的相对位置。
Unmanned aerial vehicles (UAVs) with the potential of providing reliable high-rate connectivity, are becoming a promising component of future wireless networks. A UAV collects data from a set of randomly distributed sensors, where both the locations of these sensors and their data volume to be transmitted are unknown to the UAV. In order to assist the UAV in finding the optimal motion trajectory in the face of the uncertainty without the above knowledge whilst aiming for maximizing the cumulative collected data, we formulate a reinforcement learning problem by modelling the motion-trajectory as a Markov decision process with the UAV acting as the learning agent. Then, we propose a pair of novel trajectory optimization algorithms based on stochastic modelling and reinforcement learning, which allows the UAV to optimize its flight trajectory without the need for system identification. More specifically, by dividing the considered region into small tiles, we conceive state-action-reward-state-action (Sarsa) and Q-learning based UAV-trajectory optimization algorithms (i.e., SUTOA and QUTOA) aiming to maximize the cumulative data collected during the finite flight-time. Our simulation results demonstrate that both of the proposed approaches are capable of finding an optimal trajectory under the fight-time constraint. The preference for QUTOA vs. SUTOA depends on the relative position of the start and the end points of the UAVs.