Dynamic Shielding for Reinforcement Learning in Black-Box Environments

Dynamic Shielding for Reinforcement Learning in Black-Box Environments
复制标题

黑盒环境中强化学习的动态屏蔽

DOI:
10.1007/978-3-031-19992-9_2
复制
发表时间:
2022
期刊:
Lecture Notes in Computer Science
影响因子:
--
通讯作者:
Hasuo Ichiro
Hasuo Ichiro
中科院分区:
--
文献类型:
--
作者:
Waga Masaki;Castellano Ezequiel;Pruekprasert Sasinee;Klikovits Stefan;Takisaka Toru;Hasuo Ichiro

文献摘要

参考文献

相似文献

由于学习过程中缺乏安全保证,在网络物理系统中使用强化学习(RL)具有挑战性。尽管已经有各种建议来减少学习过程中的不良行为,但这些技术大多数都需要先验的系统知识,并且它们的适用性是有限的。本文旨在减少学习过程中的不良行为,而无需任何先验的系统知识。我们提出动态屏蔽:基于模型的安全强化学习技术的扩展,称为使用自动学习的屏蔽。动态屏蔽技术使用 RPNI 算法的变体与 RL 并行构建近似系统模型,并抑制由于根据学习模型构建的屏蔽而导致的不需要的探索。通过这种组合,可以在代理经历潜在的不安全行为之前预见到这些行为。实验表明,我们的动态防护罩显着减少了训练期间意外事件的数量。
It is challenging to use reinforcement learning (RL) in cyber-physical systems due to the lack of safety guarantees during learning. Although there have been various proposals to reduce undesired behaviors during learning, most of these techniques require prior system knowledge, and their applicability is limited. This paper aims to reduce undesired behaviors during learning without requiringanyprior system knowledge. We proposedynamic shielding: an extension of a model-based safe RL technique calledshieldingusingautomata learning. The dynamic shielding technique constructs an approximate system model in parallel with RL using a variant of the RPNI algorithm and suppresses undesired explorations due to the shield constructed from the learned model. Through this combination, potentially unsafe actions can be foreseen before the agent experiences them. Experiments show that our dynamic shield significantly decreases the number of undesired events during training.
DOI: 10.1007/11901914_11
发表时间: 2006
期刊: --
影响因子: --
作者:
O. Kupferman;Robby Lampert
通讯作者: Robby Lampert
DOI: --
发表时间: 2019-04
期刊: ArXiv
影响因子: --
作者:
Maxime Bouton;J. Karlsson;A. Nakhaei;K. Fujimura;Mykel J. Kochenderfer;Jana Tumova
通讯作者: Maxime Bouton;J. Karlsson;A. Nakhaei;K. Fujimura;Mykel J. Kochenderfer;Jana Tumova
多智能体系统最低成本防护罩的综合
DOI: --
发表时间: 2019
期刊: American Control Conference
影响因子: --
作者:
Suda Bharadwaj;R. Bloem;Rayna Dimitrova;Bettina Könighofer;U. Topcu
通讯作者: U. Topcu
DOI: 10.1007/978-3-662-48395-4_4
发表时间: 2016
期刊: --
影响因子: --
作者:
Damián López;P. García
通讯作者: P. García
Shield Synthesis:反应式系统的运行时执行
DOI: --
发表时间: 2015
期刊: International Conference on Tools and Algorithms for Construction and Analysis of Systems
影响因子: --
作者:
R. Bloem;Bettina Könighofer;Robert Könighofer;Chao Wang
通讯作者: Chao Wang