Safe Reinforcement Learning Benchmark Environments for Aerospace Control Systems

Safe Reinforcement Learning Benchmark Environments for Aerospace Control Systems
复制标题

航空航天控制系统的安全强化学习基准环境

DOI:
10.1109/aero53065.2022.9843750
复制
发表时间:
2022
期刊:
IEEE Aerospace Conference
影响因子:
--
通讯作者:
Kerianne L. Hobbs
Kerianne L. Hobbs
中科院分区:
--
文献类型:
--
作者:
Umberto Ravaioli;James Cunningham;John McCarroll;Vardaan Gangal;Kyle Dunlap;Kerianne L. Hobbs

文献摘要

被引文献

相似文献

强化学习技术的最新进展证明了在高维状态空间和复杂的实时战略游戏中做出决策的能力。与以大数据集为特征的监督学习相比,用于训练强化学习代理的现有环境相对较少。此外,奖励或行动空间的微小差异会极大地改变训练环境的难度和结果。benchmark试图通过创建通用环境来解决这两个挑战,以“健身房”的形式来训练和比较强化学习技术、方法和算法。许多健身房,如经典控制和雅达利游戏环境,已经成为强化学习新研究的标准。研究人员可以很容易地在这些通用基线上比较和基准竞争的解决方案,从而实现快速创新和协作。然而,目前还没有针对航空航天问题的标准环境集,并且文献中的许多健身房不包括安全约束或运行时保证系统,当强化学习代理违反安全约束时进行干预。这份手稿描述了航空航天SafeRL框架和伴随的航空航天SafeRL基准的开发,包括交互式环境、安全约束、用于运行时保证安全监视器的软件接口和基本实现,以及一组初始基线解决方案。这一初始场景集介绍了简单的RL环境,这些环境暴露了2D和3D空气和空间问题中遇到的各种运动模式、动态和安全约束。本文还描述了这些环境的标准化评估度量,以提供与航空航天相关的一致的性能度量。这些基准测试为未来的强化学习算法、运行时保证设计和航空航天领域的神经网络验证技术提供了结构化的基础。
Recent advancements in reinforcement learning techniques demonstrate an ability to make decisions in high dimen-sional state spaces and complex real-time strategy games. In contrast to supervised learning which features large data sets, there are relatively few existing environments for training rein-forcement learning agents. In addition, small differences in re-wards or action spaces can drastically change the difficulty and results of the training environments. Benchmarks seek to tackle both of these challenges by creating common environments, in the form of “Gyms” to train and compare reinforcement learning techniques, approaches, and algorithms. Many gyms, such as the classical control and Atari games environments, have become standard in new research on reinforcement learning. Researchers can easily compare and benchmark competing so-lutions across publications on these universal baselines enabling rapid innovation and collaboration. However, there are currently no standard set of environments for aerospace problems, and many of the gyms in the literature do not include safety con-straints or run time assurance systems that intervene when the reinforcement learning agent violates safety constraints. This manuscript describes the development of the Aerospace SafeRL Framework and accompanying Aerospace SafeRL Benchmarks that include interactive environments, safety constraints, soft-ware interfaces for run time assurance safety monitors with base implementations, and an initial set of baseline solutions. This initial set of scenarios introduces simple RL environments that expose the kinds of motion patterns, dynamics, and safety constraints encountered in air and space problems in 2D and 3D. This manuscript also describes standardized evaluation metrics for these environments to provide a consistent performance measurement with aerospace relevance. These benchmarks pro-vide a structured foundation for future reinforcement learning algorithms, run time assurance designs, and neural network verification techniques for the aerospace domain.