Interestingness Elements for Explainable Reinforcement Learning: Understanding Agents' Capabilities and Limitations

Interestingness Elements for Explainable Reinforcement Learning: Understanding Agents' Capabilities and Limitations
复制标题

DOI:
10.1016/j.artint.2020.103367
复制
发表时间:
2019-12
期刊:
ArXiv
影响因子:
--
通讯作者:
P. Sequeira;M. Gervasio
P. Sequeira;M. Gervasio
中科院分区:
其他
文献类型:
--
作者:
P. Sequeira;M. Gervasio

文献摘要

被引文献

相似文献

我们提出了一个可解释的强化学习(XRL)框架,该框架分析了代理与环境交互的历史,以提取有助于解释其行为的兴趣元素。该框架依赖于从标准RL算法中容易获得的数据,这些数据可以很容易地由代理在学习时收集。我们描述了如何创建一个代理的行为在短视频剪辑的形式突出关键的互动时刻,建议的元素的基础上的视觉摘要。我们还报告了一项用户研究,在该研究中,我们评估了人类正确感知具有不同特征的代理人的能力,包括他们的能力和局限性,并给出了我们的框架自动生成的视觉摘要。结果表明,不同兴趣度元素所捕获的方面的多样性对于帮助人类正确理解智能体在执行任务时的优势和局限性至关重要,并确定何时可能需要调整以提高其性能。
We propose an explainable reinforcement learning (XRL) framework that analyzes an agent's history of interaction with the environment to extract interestingness elements that help explain its behavior. The framework relies on data readily available from standard RL algorithms, augmented with data that can easily be collected by the agent while learning. We describe how to create visual summaries of an agent's behavior in the form of short video-clips highlighting key interaction moments, based on the proposed elements. We also report on a user study where we evaluated the ability of humans to correctly perceive the aptitude of agents with different characteristics, including their capabilities and limitations, given visual summaries automatically generated by our framework. The results show that the diversity of aspects captured by the different interestingness elements is crucial to help humans correctly understand an agent's strengths and limitations in performing a task, and determine when it might need adjustments to improve its performance.