Explicable Policy Search

Explicable Policy Search
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Ze Gong
Ze Gong
中科院分区:
其他
文献类型:
--
作者:
Ze Gong

文献摘要

被引文献

相似文献

人类队友在互动过程中经常会对彼此形成有意识和潜意识的期望。团队合作的成功取决于这些期望能否得到满足。类似地,对于一个智能体在人类旁边工作,它必须考虑人类对其行为的期望。忽视这些期望会导致信任丧失和团队绩效下降。这里的一个关键挑战是人类的期望可能与智能体的最优行为不一致,例如,由于人类对任务领域的部分或不准确理解。先前关于可解释规划的工作描述了智能体通过在任务绩效和更符合预期或“可解释”的行为之间进行权衡来尊重其人类队友期望的能力。在本文中,我们引入可解释策略搜索(EPS),以便在具有连续状态和动作空间的强化学习(RL)环境中将这种能力显著扩展到随机领域。此外,与传统的RL方法不同,EPS必须同时推断人类隐藏的期望。这种推断需要关于人类对领域动态的信念及其奖励模型的信息,但直接询问这些是不切实际的。我们证明,这种信息可以通过EPS的一个替代奖励函数进行必要且充分的编码,该函数可以基于人类对智能体行为的反馈来学习。然后,该替代奖励函数被用于重塑智能体的奖励函数,这被证明等同于搜索一个可解释的策略。我们在一组具有合成人类模型的导航领域以及一个通过用户研究的自动驾驶领域中评估EPS。结果表明,我们的方法能够生成可解释的行为,这些行为能够智能地协调任务绩效和人类期望,并且在人 - 智能体团队合作领域具有现实相关性。
Human teammates often form conscious and subconscious expectations of each other during interaction. Teaming success is contingent on whether such expectations can be met. Similarly, for an intelligent agent to operate beside a human, it must consider the human’s expectation of its behavior. Disregarding such expectations can lead to the loss of trust and degraded team performance. A key challenge here is that the human’s expectation may not align with the agent’s optimal behavior, e.g., due to the human’s partial or inaccurate understanding of the task domain. Prior work on explicable planning described the ability of agents to respect their human teammate’s expectations by trading off task performance for more expected or “ explicable ” behaviors. In this paper, we introduce Explicable Policy Search (EPS) to significantly extend such an ability to stochastic domains in a reinforcement learning (RL) setting with continuous state and action spaces. Furthermore, in contrast to the traditional RL methods, EPS must at the same time infer the human’s hidden expectations. Such inferences require information about the human’s belief about the domain dynamics and her reward model but directly querying them is impractical. We demonstrate that such information can be necessarily and sufficiently encoded by a surrogate reward function for EPS, which can be learned based on the human’s feedback on the agent’s behavior. The surrogate reward function is then used to reshape the agent’s reward function, which is shown to be equivalent to searching for an explicable policy. We evaluate EPS in a set of navigation domains with synthetic human models and in an autonomous driving domain with a user study. The results suggest that our method can generate explicable behaviors that reconcile task performance with human expectations intelligently and has real-world relevance in human-agent teaming domains.