Autonomous Drone Racing with Deep Reinforcement Learning

Autonomous Drone Racing with Deep Reinforcement Learning
复制标题

DOI:
10.1109/iros51168.2021.9636053
复制
发表时间:
2021-03
期刊:
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Yunlong Song;Mats Steinweg;Elia Kaufmann;D. Scaramuzza
Yunlong Song;Mats Steinweg;Elia Kaufmann;D. Scaramuzza
中科院分区:
其他
文献类型:
--
作者:
Yunlong Song;Mats Steinweg;Elia Kaufmann;D. Scaramuzza

文献摘要

被引文献

相似文献

在许多机器人任务中,例如自主无人机竞速,目标是尽可能快地穿过一系列航路点。这项任务的一个关键挑战是规划时间最优轨迹,这通常是通过预先假定完全了解要经过的航路点来解决的。由此产生的解决方案要么高度针对单一赛道布局进行了专门设计,要么由于对平台动力学进行了简化假设而并非最优。在这项工作中,提出了一种用于四旋翼飞行器生成近时间最优轨迹的新方法。利用深度强化学习和相对门观测,我们的方法能够计算近时间最优轨迹,并使轨迹适应环境变化。对于非平凡的赛道配置,我们的方法相较于基于轨迹优化的方法具有计算优势。所提出的方法在模拟和现实世界的一组赛道上进行了评估,使用一架实体四旋翼飞行器达到了高达60千米/小时的速度。
In many robotic tasks, such as autonomous drone racing, the goal is to travel through a set of waypoints as fast as possible. A key challenge for this task is planning the timeoptimal trajectory, which is typically solved by assuming perfect knowledge of the waypoints to pass in advance. The resulting solution is either highly specialized for a single-track layout, or suboptimal due to simplifying assumptions about the platform dynamics. In this work, a new approach to near-time-optimal trajectory generation for quadrotors is presented. Leveraging deep reinforcement learning and relative gate observations, our approach can compute near-time-optimal trajectories and adapt the trajectory to environment changes. Our method exhibits computational advantages over approaches based on trajectory optimization for non-trivial track configurations. The proposed approach is evaluated on a set of race tracks in simulation and the real world, achieving speeds of up to 60kmh−1 with a physical quadrotor.