Safety Verification of Cyber-Physical Systems with Reinforcement Learning Control

Safety Verification of Cyber-Physical Systems with Reinforcement Learning Control
复制标题

DOI:
10.1145/3358230
复制
发表时间:
2019-10-01
影响因子:
2
通讯作者:
Koutsoukos, Xenofon
Koutsoukos, Xenofon
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hoang-Dung Tran;Cai, Feiyang;Koutsoukos, Xenofon

文献摘要

被引文献

相似文献

本文提出了一种新的远程达到性分析方法,以通过增强学习控制器验证网络物理系统(CPS)的安全性。我们方法的基础在于使用Star Sets对神经网络控制系统的两种有效,精确和过度陈列的可及性算法,这是Polyhedra的有效表示。使用这些算法,我们通过逐步搜索无法确定系统安全性的关键初始条件来确定具有神经网络控制器安全至关重要系统的初始条件。我们的方法会产生严格的过度评估误差,并且在计算上是有效的,这使应用程序可以使用学习启用组件(LEC)实用的CPS。我们在NNV中实施方法,NNV是一种用于神经网络和神经网络控制系统的验证工具,并通过使用增强型学习器(RL)控制器验证实用的高级紧急制动系统(AEB)来评估其优势和适用性深层确定性策略梯度(DDPG)方法。实验结果表明,我们的新可达性算法比现有的基于多面体的方法保守得多。我们成功地通过RL控制器确定了AEB的初始条件的整个区域,以确保系统的安全性,而基于多面体的方法无法证明系统的安全性。
This paper proposes a new forward reachability analysis approach to verify safety of cyber-physical systems (CPS) with reinforcement learning controllers. The foundation of our approach lies on two efficient, exact and over-approximate reachability algorithms for neural network control systems using star sets, which is an efficient representation of polyhedra. Using these algorithms, we determine the initial conditions for which a safety-critical system with a neural network controller is safe by incrementally searching a critical initial condition where the safety of the system cannot be established. Our approach produces tight over-approximation error and it is computationally efficient, which allows the application to practical CPS with learning enable components (LECs). We implement our approach in NNV, a recent verification tool for neural networks and neural network control systems, and evaluate its advantages and applicability by verifying safety of a practical Advanced Emergency Braking System (AEBS) with a reinforcement learning (RL) controller trained using the deep deterministic policy gradient (DDPG) method. The experimental results show that our new reachability algorithms are much less conservative than existing polyhedra-based approaches. We successfully determine the entire region of the initial conditions of the AEBS with the RL controller such that the safety of the system is guaranteed, while a polyhedra-based approach cannot prove the safety properties of the system.