Actuator Trajectory Planning for UAVs with Overhead Manipulator using Reinforcement Learning

Actuator Trajectory Planning for UAVs with Overhead Manipulator using Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2308.12843
复制
发表时间:
2023-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Hazim Alzorgan;Abolfazl Razi;A. Moshayedi
Hazim Alzorgan;Abolfazl Razi;A. Moshayedi
中科院分区:
其他
文献类型:
--
作者:
Hazim Alzorgan;Abolfazl Razi;A. Moshayedi

文献摘要

相似文献

本文研究了一种空中机械臂系统的操作,即一种配备有两个自由度的可控臂的无人机(UAV),用于在飞行中执行驱动任务。我们的解决方案是基于采用Q-learning方法来控制手臂尖端的轨迹,也称为末端执行器。更具体地说,我们开发了一个基于碰撞时间(TTC)的运动规划模型,该模型使四旋翼无人机能够在确保机械手可达性的同时绕过障碍物。此外,我们利用基于模型的q -学习模型来独立跟踪和控制机械臂末端执行器的期望轨迹,给定无人机平台的任意基线轨迹。这样的组合可以实现各种驱动任务,例如高空焊接,结构监测和维修,电池更换,排水沟清洁,摩天大楼清洁以及在难以到达和危险环境中的电源线维护,同时保持与飞行控制固件的兼容性。我们基于rl的控制机制产生了一种鲁棒控制策略,可以处理无人机运动中的不确定性,提供了很好的性能。具体来说,我们的方法在平均位移误差(即目标与获得的轨迹点之间的平均距离)方面实现了92%的准确率,使用了15,000集的Q-learning
In this paper, we investigate the operation of an aerial manipulator system, namely an Unmanned Aerial Vehicle (UAV) equipped with a controllable arm with two degrees of freedom to carry out actuation tasks on the fly. Our solution is based on employing a Q-learning method to control the trajectory of the tip of the arm, also called end-effector. More specifically, we develop a motion planning model based on Time To Collision (TTC), which enables a quadrotor UAV to navigate around obstacles while ensuring the manipulator's reachability. Additionally, we utilize a model-based Q-learning model to independently track and control the desired trajectory of the manipulator's end-effector, given an arbitrary baseline trajectory for the UAV platform. Such a combination enables a variety of actuation tasks such as high-altitude welding, structural monitoring and repair, battery replacement, gutter cleaning, skyscrapper cleaning, and power line maintenance in hard-to-reach and risky environments while retaining compatibility with flight control firmware. Our RL-based control mechanism results in a robust control strategy that can handle uncertainties in the motion of the UAV, offering promising performance. Specifically, our method achieves 92% accuracy in terms of average displacement error (i.e. the mean distance between the target and obtained trajectory points) using Q-learning with 15,000 episodes