Learning to Act Safely With Limited Exposure and Almost Sure Certainty

Learning to Act Safely With Limited Exposure and Almost Sure Certainty
复制标题

DOI:
10.1109/tac.2023.3240925
复制
发表时间:
2021-05
影响因子:
6.8
通讯作者:
Agustin Castellano;Hancheng Min;J. Bazerque;Enrique Mallada
Agustin Castellano;Hancheng Min;J. Bazerque;Enrique Mallada
中科院分区:
计算机科学2区
文献类型:
--
作者:
Agustin Castellano;Hancheng Min;J. Bazerque;Enrique Mallada

文献摘要

被引文献

相似文献

本文提出了这样一个概念,即学习在未知环境中采取安全行动,即使概率为1,也可以实现,而不需要无限数量的探索性试验。这确实是可能的,只要人们愿意在最优性、暴露于不安全事件的水平和不安全操作的最大检测时间之间进行权衡。我们说明了这一概念在两个互补的设置。首先,我们专注于规范的多臂强盗问题,并研究在存在不确定性的学习安全的内在权衡。在充分探索的温和假设下,我们提供了一个算法,可证明检测所有不安全的机器在(预期)有限数量的回合。该分析还揭示了保护环境所需的轮数与丢弃安全机器的概率之间的权衡。然后,我们考虑的问题,找到最佳的政策马尔可夫决策过程(MDP)几乎肯定的约束。我们表明,行动价值函数满足基于障碍的分解,允许识别独立的奖励过程的可行的政策。使用这种分解,我们开发了一个障碍学习算法,识别这种不安全的状态动作对在一个有限的预期步骤数。我们的分析进一步强调了底层MDP检测不安全操作所需的时间延迟与暴露于不安全事件的水平之间的权衡。仿真证实了我们的理论研究结果,进一步说明了上述权衡,并表明安全约束可以加快学习过程。
This article puts forward the concept that learning to take safe actions in unknown environments, even with probability one guarantees, can be achieved without the need for an unbounded number of exploratory trials. This is indeed possible, provided that one is willing to navigate tradeoffs between optimality, level of exposure to unsafe events, and the maximum detection time of unsafe actions. We illustrate this concept in two complementary settings. We first focus on the canonical multiarmed bandit problem and study the intrinsic tradeoffs of learning safety in the presence of uncertainty. Under mild assumptions on sufficient exploration, we provide an algorithm that provably detects all unsafe machines in an (expected) finite number of rounds. The analysis also unveils a tradeoff between the number of rounds needed to secure the environment and the probability of discarding safe machines. We then consider the problem of finding optimal policies for a Markov decision process (MDP) with almost sure constraints. We show that the action-value function satisfies a barrier-based decomposition that allows for the identification of feasible policies independently of the reward process. Using this decomposition, we develop a barrier-learning algorithm, that identifies such unsafe state–action pairs in a finite expected number of steps. Our analysis further highlights a tradeoff between the time lag for the underlying MDP necessary to detect unsafe actions, and the level of exposure to unsafe events. Simulations corroborate our theoretical findings, further illustrating the aforementioned tradeoffs, and suggesting that safety constraints can speed up the learning process.