Shield Synthesis for Reinforcement Learning

Shield Synthesis for Reinforcement Learning
复制标题

强化学习的盾牌合成

DOI:
10.1007/978-3-030-61362-4_16
复制
发表时间:
2020
期刊:
Proceedings of the Institution of Mechanical Engineers, Part O: Journal of Risk and Reliability
影响因子:
--
通讯作者:
R. Bloem
R. Bloem
中科院分区:
--
文献类型:
--
作者:
Bettina Könighofer;Florian Lorber;N. Jansen;R. Bloem

文献摘要

参考文献

被引文献

相似文献

强化学习算法发现最大化奖励的策略。然而,这些政策通常不坚持安全,使得强化学习(以及整个人工智能)的安全成为一个悬而未决的研究问题。屏蔽合成是一种正式的方法,用于合成被称为屏蔽的按结构正确的反应系统,该系统在确保运行系统的安全属性的同时,尽可能减少对其运行的干扰。附着在学习代理上的盾牌确保了学习和执行阶段的安全。本文总结了由不同规格说明语言合成的三种屏蔽,并讨论了它们在强化学习中的适用性。首先,我们讨论执行表示为线性时态逻辑规范的规范的确定性屏蔽。其次,我们讨论了概率时序逻辑中由规格说明合成概率屏蔽的问题。第三,我们讨论了如何从时间自动机的规范中综合出时间屏蔽。本文综述了这三种屏蔽材料的应用领域、优缺点和合成方法,并对实验结果进行了综述。
Reinforcement learning algorithms discover policies that maximize reward. However, these policies generally do not adhere to safety, leaving safety in reinforcement learning (and in artificial intelligence in general) an open research problem. Shield synthesis is a formal approach to synthesize a correct-by-construction reactive system called a shield that enforces safety properties of a running system while interfering with its operation as little as possible. A shield attached to a learning agent guarantees safety during learning and execution phases. In this paper we summarize three types of shields that are synthesized from different specification languages, and discuss their applicability to reinforcement learning. First, we discuss deterministic shields that enforce specifications expressed as linear temporal logic specifications. Second, we discuss the synthesis of probabilistic shields from specifications in probabilistic temporal logic. Third, we discuss how to synthesize timed shields from timed automata specifications. This paper summarizes the application areas, advantages, disadvantages and synthesis approaches for the three types of shields and gives an overview of experimental results.
护盾合成
DOI: 10.1007/978-3-319-49052-6_9
发表时间: 2017
影响因子: 0.8
作者:
Alshiekh, Mohammed;Bloem, Roderick;Humphrey, Laura;Topcu, Ufuk;Wang, Chao
通讯作者: Wang, Chao