Safety-Aware Apprenticeship Learning

Safety-Aware Apprenticeship Learning
复制标题

DOI:
10.1007/978-3-319-96145-3_38
复制
发表时间:
2017-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Weichao Zhou;Wenchao Li
Weichao Zhou;Wenchao Li
中科院分区:
其他
文献类型:
--
作者:
Weichao Zhou;Wenchao Li

文献摘要

被引文献

相似文献

学徒学习(AL)是一种从演示中学习的技术,其中马尔可夫决策过程(MDP)的奖励函数对于学习代理来说是未知的,并且代理必须通过观察专家的演示来得出良好的策略。在本文中,我们研究了如何使 AL 算法本质上安全,同时仍然满足其学习目标的问题。我们考虑这样一种设置,其中未知奖励函数被假设为一组状态特征的线性组合,并且安全属性在概率计算树逻辑(PCTL)中指定。通过在 AL 中嵌入概率模型检查,我们提出了一种新颖的反例引导方法,可以确保安全性,同时保留学习策略的性能。我们在安全至关重要的几个具有挑战性的 AL 场景中展示了我们的方法的有效性。
Apprenticeship learning (AL) is a kind of Learning from Demonstration techniques where the reward function of a Markov Decision Process (MDP) is unknown to the learning agent and the agent has to derive a good policy by observing an expert’s demonstrations. In this paper, we study the problem of how to make AL algorithms inherently safe while still meeting its learning objective. We consider a setting where the unknown reward function is assumed to be a linear combination of a set of state features, and the safety property is specified in Probabilistic Computation Tree Logic (PCTL). By embedding probabilistic model checking inside AL, we propose a novelcounterexample-guidedapproach that can ensure safety while retaining performance of the learnt policy. We demonstrate the effectiveness of our approach on several challenging AL scenarios where safety is essential.