Learning to Fly—a Gym Environment with PyBullet Physics for Reinforcement Learning of Multi-agent Quadcopter Control

Learning to Fly—a Gym Environment with PyBullet Physics for Reinforcement Learning of Multi-agent Quadcopter Control
复制标题

DOI:
10.1109/iros51168.2021.9635857
复制
发表时间:
2021-03
期刊:
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Jacopo Panerati;Hehui Zheng;Siqi Zhou;James Xu;Amanda Prorok;Angela P. Schoellig University of Toronto Institute for A Studies-Angela-P.-Schoellig-University-of-Toronto-Institute-2051868118;Vector Institute for Artificial Intelligence;U. Cambridge
Jacopo Panerati;Hehui Zheng;Siqi Zhou;James Xu;Amanda Prorok;Angela P. Schoellig University of Toronto Institute for A Studies-Angela-P.-Schoellig-University-of-Toronto-Institute-2051868118;Vector Institute for Artificial Intelligence;U. Cambridge
中科院分区:
其他
文献类型:
--
作者:
Jacopo Panerati;Hehui Zheng;Siqi Zhou;James Xu;Amanda Prorok;Angela P. Schoellig University of Toronto Institute for A Studies-Angela-P.-Schoellig-University-of-Toronto-Institute-2051868118;Vector Institute for Artificial Intelligence;U. Cambridge

文献摘要

被引文献

相似文献

机器人模拟器对于学术研究和教育以及安全关键应用的开发至关重要。强化学习环境-简单的模拟加上奖励函数形式的问题规范-对于标准化学习算法的开发(和基准测试)也很重要。然而,全尺寸模拟器通常缺乏可移植性和可扩展性。反之亦然,许多强化学习环境在玩具般的问题中权衡现实主义和高样本吞吐量。虽然公共数据集极大地促进了深度学习和计算机视觉,但我们仍然缺乏软件工具来同时开发控制理论和强化学习方法。在本文中,我们基于Bullet物理引擎为多个四轴飞行器提出了一个开源的OpenAI Gym-like环境。它的多智能体和基于视觉的强化学习界面,以及对现实碰撞和空气动力学效果的支持,使其成为同类产品中的第一个。我们通过几个例子展示了它的使用,无论是控制(轨迹跟踪与PID控制,多机器人飞行与下洗,等)。或强化学习(单智能体和多智能体稳定任务),希望能激发未来结合控制理论和机器学习的研究。
Robotic simulators are crucial for academic research and education as well as the development of safety-critical applications. Reinforcement learning environments— simple simulations coupled with a problem specification in the form of a reward function—are also important to standardize the development (and benchmarking) of learning algorithms. Yet, full-scale simulators typically lack portability and paral-lelizability. Vice versa, many reinforcement learning environments trade-off realism for high sample throughputs in toy-like problems. While public data sets have greatly benefited deep learning and computer vision, we still lack the software tools to simultaneously develop—and fairly compare—control theory and reinforcement learning approaches. In this paper, we propose an open-source OpenAI Gym-like environment for multiple quadcopters based on the Bullet physics engine. Its multi-agent and vision-based reinforcement learning interfaces, as well as the support of realistic collisions and aerodynamic effects, make it, to the best of our knowledge, a first of its kind. We demonstrate its use through several examples, either for control (trajectory tracking with PID control, multi-robot flight with downwash, etc.) or reinforcement learning (single and multi-agent stabilization tasks), hoping to inspire future research that combines control theory and machine learning.