Verifiable RNN-Based Policies for POMDPs Under Temporal Logic Constraints

Verifiable RNN-Based Policies for POMDPs Under Temporal Logic Constraints
复制标题

时态逻辑约束下可验证的基于 RNN 的 POMDP 策略

DOI:
--
复制
发表时间:
2020
期刊:
International Joint Conference on Artificial Intelligence
影响因子:
--
通讯作者:
U. Topcu
U. Topcu
中科院分区:
--
文献类型:
--
作者:
Steven Carr;N. Jansen;U. Topcu

文献摘要

参考文献

被引文献

相似文献

递归神经网络(RNN)已经成为顺序决策问题中控制策略的有效表示。 然而,基于RNN的策略应用的一个主要缺点是难以提供对行为规范(例如安全性和/或可达性)的满足的正式保证。 通过整合形式化方法和机器学习的技术,我们提出了一种从RNN中自动提取有限状态控制器(FSC)的方法,当RNN与有限状态系统模型组合时,可以使用现有的形式化验证工具。 具体来说,我们引入了一个迭代修改所谓的量化瓶颈插入技术,以创建一个FSC作为一个随机化的策略与内存。 对于所得到的FSC不满足规范的情况,验证生成诊断信息。 我们利用这些信息来调整提取的FSC中的内存量,或者对RNN进行集中的重新训练。 虽然普遍适用,我们详细说明了部分可观察马尔可夫决策过程(POMDPs),这是众所周知的是非常困难的政策合成的背景下,由此产生的迭代过程。 数值实验表明,在最优基准值的2%以内,该方法比传统的POMDP综合方法的性能提高了3个数量级。
Recurrent neural networks (RNNs) have emerged as an effective representation of control policies in sequential decision-making problems. However, a major drawback in the application of RNN-based policies is the difficulty in providing formal guarantees on the satisfaction of behavioral specifications, e.g. safety and/or reachability. By integrating techniques from formal methods and machine learning, we propose an approach to automatically extract a finite-state controller (FSC) from an RNN, which, when composed with a finite-state system model, is amenable to existing formal verification tools. Specifically, we introduce an iterative modification to the so-called quantized bottleneck insertion technique to create an FSC as a randomized policy with memory. For the cases in which the resulting FSC fails to satisfy the specification, verification generates diagnostic information. We utilize this information to either adjust the amount of memory in the extracted FSC or perform focused retraining of the RNN. While generally applicable, we detail the resulting iterative procedure in the context of policy synthesis for partially observable Markov decision processes (POMDPs), which is known to be notoriously hard. The numerical experiments show that the proposed approach outperforms traditional POMDP synthesis methods by 3 orders of magnitude within 2% of optimal benchmark values.
DOI: 10.1007/s11241-017-9269-4
发表时间: 2017-05-01
期刊: REAL-TIME SYSTEMS
影响因子: 1.3
作者:
Norman, Gethin;Parker, David;Zou, Xueyi
通讯作者: Zou, Xueyi