Failing with Grace: Learning Neural Network Controllers that are Boundedly Unsafe

Failing with Grace: Learning Neural Network Controllers that are Boundedly Unsafe
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Panagiotis Vlantis;M. Zavlanos
Panagiotis Vlantis;M. Zavlanos
中科院分区:
其他
文献类型:
--
作者:
Panagiotis Vlantis;M. Zavlanos

文献摘要

相似文献

在这项工作中,我们考虑的问题,学习前馈神经网络控制器,以安全地引导一个任意形状的平面机器人在一个紧凑的和障碍物闭塞的工作空间。与强烈依赖于接近安全状态空间边界的数据点的密度来训练具有闭环安全保证的神经网络控制器的现有方法不同,在这里,我们提出了一种替代方法,该方法提升了在实践中难以满足的数据上的强假设,而是允许优雅的安全违规,即,具有可以在空间上控制的有界幅度。为此,我们采用可达性分析技术来封装训练过程中的安全约束。具体来说,为了获得一个计算效率过逼近的前向可达集的闭环系统,我们划分的机器人的状态空间成细胞和自适应细分的细胞,包含的状态,可能会逃脱安全集下的训练控制律。然后,使用每个单元格的前向可达集和不可行的机器人配置集之间的重叠作为违反安全性的措施,我们引入适当的条款到损失函数中,惩罚这种重叠在训练过程中。其结果是,我们的方法可以学习一个安全的向量场的闭环系统,并在同一时间,提供最坏情况下的边界上的安全违反整个配置空间,定义的过度逼近的前向可达集的闭环系统和不安全的状态集之间的重叠。此外,它可以控制计算复杂度和这些边界的紧密性之间的权衡。我们所提出的方法是支持的理论结果和仿真研究。
In this work, we consider the problem of learning a feed-forward neural network controller to safely steer an arbitrarily shaped planar robot in a compact and obstacle-occluded workspace. Unlike existing methods that depend strongly on the density of data points close to the boundary of the safe state space to train neural network controllers with closed-loop safety guarantees, here we propose an alternative approach that lifts such strong assumptions on the data that are hard to satisfy in practice and instead allows for graceful safety violations, i.e., of a bounded magnitude that can be spatially controlled. To do so, we employ reachability analysis techniques to encapsulate safety constraints in the training process. Specifically, to obtain a computationally efficient over-approximation of the forward reachable set of the closed-loop system, we partition the robot's state space into cells and adaptively subdivide the cells that contain states which may escape the safe set under the trained control law. Then, using the overlap between each cell's forward reachable set and the set of infeasible robot configurations as a measure for safety violations, we introduce appropriate terms into the loss function that penalize this overlap in the training process. As a result, our method can learn a safe vector field for the closed-loop system and, at the same time, provide worst-case bounds on safety violation over the whole configuration space, defined by the overlap between the over-approximation of the forward reachable set of the closed-loop system and the set of unsafe states. Moreover, it can control the tradeoff between computational complexity and tightness of these bounds. Our proposed method is supported by both theoretical results and simulation studies.