Fast and Accurate Trajectory Tracking for Unmanned Aerial Vehicles based on Deep Reinforcement Learning

Fast and Accurate Trajectory Tracking for Unmanned Aerial Vehicles based on Deep Reinforcement Learning
复制标题

DOI:
10.1109/rtcsa.2019.8864571
复制
发表时间:
2019-08
期刊:
2019 IEEE 25th International Conference on Embedded and Real-Time Computing Systems and Applications (RTCSA)
影响因子:
--
通讯作者:
Yilan Li;Hongjia Li;Zhe Li;Haowen Fang;A. Sanyal;Yanzhi Wang;Qinru Qiu
Yilan Li;Hongjia Li;Zhe Li;Haowen Fang;A. Sanyal;Yanzhi Wang;Qinru Qiu
中科院分区:
其他
文献类型:
--
作者:
Yilan Li;Hongjia Li;Zhe Li;Haowen Fang;A. Sanyal;Yanzhi Wang;Qinru Qiu

文献摘要

被引文献

相似文献

固定翼无人机的连续轨迹控制是一个复杂的问题,需要考虑隐动力学的影响。由于无人机具有多自由度,基于传统控制理论的跟踪方法,如比例-积分-微分(PID)控制方法,在响应时间和调节鲁棒性方面存在局限性,而基于模型的方法,即根据无人机当前状态计算力和力矩的方法,则复杂且刚性。我们提出了一个行动者-评论家强化学习框架,通过一组所需的航路点控制无人机轨迹。构建深度神经网络来学习最优跟踪策略,并开发强化学习来优化所产生的跟踪方案。实验结果表明,我们提出的方法可以实现58.14%的位置误差,减少21.77%的系统功耗和9.23%的速度比基线。执行器网络仅由线性运算组成,因此基于现场可编程门阵列(FPGA)的硬件加速可以很容易地设计用于节能的实时控制。
Continuous trajectory control of fixed-wing unmanned aerial vehicles (UAVs) is complicated when considering hidden dynamics. Due to UAV multi degrees of freedom, tracking methodologies based on conventional control theory, such as Proportional-Integral-Derivative (PID) has limitations in response time and adjustment robustness, while a model based approach that calculates the force and torques based on UAV's current status is complicated and rigid. We present an actor-critic reinforcement learning framework that controls UAV trajectory through a set of desired waypoints. A deep neural network is constructed to learn the optimal tracking policy and reinforcement learning is developed to optimize the resulting tracking scheme. The experimental results show that our proposed approach can achieve 58.14% less position error, 21.77% less system power consumption and 9.23% faster attainment than the baseline. The actor network consists of only linear operations, hence Field Programmable Gate Arrays (FPGA) based hardware acceleration can easily be designed for energy efficient real-time control.