Learning to be safe, in finite time

Learning to be safe, in finite time
复制标题

DOI:
10.23919/acc50511.2021.9482829
复制
发表时间:
2020-10
期刊:
2021 American Control Conference (ACC)
影响因子:
--
通讯作者:
Agustin Castellano;J. Bazerque;Enrique Mallada
Agustin Castellano;J. Bazerque;Enrique Mallada
中科院分区:
其他
文献类型:
--
作者:
Agustin Castellano;J. Bazerque;Enrique Mallada

文献摘要

相似文献

本文的目的是提出的概念,学习在未知的环境中采取安全的行动,即使有概率的保证,可以实现,而不需要一个无限数量的探索性试验,只要一个愿意放松其最优性要求温和。我们专注于典型的多臂强盗问题,并寻求研究安全学习内在的探索-保存权衡。更确切地说,通过定义一个障碍度量,计数的不安全的动作的数量,我们提供了一个算法,丢弃不安全的机器(或动作),概率为1,实现恒定的障碍。我们的算法是植根于经典的序贯概率比测试,重新定义这里的连续任务。在充分探索的标准假设下,我们的规则可证明在(预期)有限数量的轮中检测到所有不安全的机器。该分析还揭示了保护环境所需的回合数与丢弃安全机器的概率之间的权衡。我们的决策规则可以围绕任何其他算法来优化特定的辅助目标,因为它提供了一个安全的环境来搜索(近似)最优策略。模拟证实了我们的理论研究结果,并进一步说明了上述权衡。
This paper aims to put forward the concept that learning to take safe actions in unknown environments, even with probability one guarantees, can be achieved without the need for an unbounded number of exploratory trials, provided that one is willing to relax its optimality requirements mildly. We focus on the canonical multi-armed bandit problem and seek to study the exploration-preservation trade-off intrinsic within safe learning. More precisely, by defining a handicap metric that counts the number of unsafe actions, we provide an algorithm for discarding unsafe machines (or actions), with probability one, that achieves constant handicap. Our algorithm is rooted in the classical sequential probability ratio test, redefined here for continuing tasks. Under standard assumptions on sufficient exploration, our rule provably detects all unsafe machines in an (expected) finite number of rounds. The analysis also unveils a trade-off between the number of rounds needed to secure the environment and the probability of discarding safe machines. Our decision rule can wrap around any other algorithm to optimize a specific auxiliary goal since it provides a safe environment to search for (approximately) optimal policies. Simulations corroborate our theoretical findings and further illustrate the aforementioned trade-offs.