Safe-Control-Gym: A Unified Benchmark Suite for Safe Learning-Based Control and Reinforcement Learning in Robotics

Safe-Control-Gym: A Unified Benchmark Suite for Safe Learning-Based Control and Reinforcement Learning in Robotics
复制标题

DOI:
10.1109/lra.2022.3196132
复制
发表时间:
2021-09
影响因子:
5.2
通讯作者:
Zhaocong Yuan;Adam W. Hall;Siqi Zhou;Lukas Brunke;Melissa Greeff;Jacopo Panerati;Angela P. Schoellig
Zhaocong Yuan;Adam W. Hall;Siqi Zhou;Lukas Brunke;Melissa Greeff;Jacopo Panerati;Angela P. Schoellig
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhaocong Yuan;Adam W. Hall;Siqi Zhou;Lukas Brunke;Melissa Greeff;Jacopo Panerati;Angela P. Schoellig

文献摘要

被引文献

相似文献

近年来,强化学习和基于学习的控制,以及对它们的安全性的研究,这对于在现实世界中部署机器人至关重要,都获得了巨大的吸引力。然而,为了充分衡量新结果的进展和适用性,我们需要工具来公平地比较控制和强化学习社区提出的方法。在这里,我们提出了一个新的开源基准套件,称为安全控制健身房,支持基于模型和基于数据的控制技术。我们提供了三个动态系统的实现-车杆,一维和二维四旋翼和两个控制任务-稳定和轨迹跟踪。我们建议扩展OpenAI的Gym API-强化学习研究中的事实标准-(i)指定(和查询)符号动力学的能力和(ii)约束,以及(iii)(可重复)在控制输入,状态测量和惯性属性中注入模拟干扰。为了证明我们的建议,并试图使研究社区更紧密地联系在一起,我们展示了如何使用safe-control-gym来定量比较传统控制,基于学习的控制和强化学习领域的多种方法的控制性能,数据效率和安全性。
In recent years, both reinforcement learning and learning-based control—as well as the study of their safety, which is crucial for deployment in real-world robots—have gained significant traction. However, to adequately gauge the progress and applicability of new results, we need the tools to equitably compare the approaches proposed by the controls and reinforcement learning communities. Here, we propose a new open-source benchmark suite, called safe-control-gym, supporting both model-based and data-based control techniques. We provide implementations for three dynamic systems—the cart-pole, the 1D, and 2D quadrotor—and two control tasks—stabilization and trajectory tracking. We propose to extend OpenAI's Gym API—the de facto standard in reinforcement learning research—with (i) the ability to specify (and query) symbolic dynamics and (ii) constraints, and (iii) (repeatably) inject simulated disturbances in the control inputs, state measurements, and inertial properties. To demonstrate our proposal and in an attempt to bring research communities closer together, we show how to use safe-control-gym to quantitatively compare the control performance, data efficiency, and safety of multiple approaches from the fields of traditional control, learning-based control, and reinforcement learning.