Explainable Artificial Intelligence - First World Conference, xAI 2023, Lisbon, Portugal, July 26-28, 2023, Proceedings, Part II

Explainable Artificial Intelligence - First World Conference, xAI 2023, Lisbon, Portugal, July 26-28, 2023, Proceedings, Part II
复制标题

可解释的人工智能 - 第一届世界会议,xAI 2023,葡萄牙里斯本,2023 年 7 月 26-28 日,会议记录,第二部分

DOI:
10.1007/978-3-031-44067-0_4
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Liu X
Liu X
中科院分区:
--
文献类型:
--
作者:
Liu X

文献摘要

相似文献

在反事实推理的辅助下,因果归因被认为是人类解释的一个关键特征。在本文中,我们提出了一个事后对比解释框架,强化学习(RL)的基础上比较学习政策下的实际环境奖励与假设(反事实)奖励。该框架通过访问学习的Q函数和识别交叉的临界状态来提供政策层面的解释。全局解释是通过基于这些状态的子轨迹的可视化来总结政策行为,而局部解释是基于状态中的动作值。我们在几个网格世界的例子进行实验。我们的研究结果表明,这是可能的,以解释基于Q函数的学习策略之间的差异。这证明了在部署策略时更明智的人类决策的潜力,并强调了在RL中开发进一步XAI技术的可能性。
Causal attribution aided by counterfactual reasoning is recognised as a key feature of human explanation. In this paper we propose a post-hoc contrastive explanation framework for reinforcement learning (RL) based on comparing learned policies under actual environmental rewards vs. hypothetical (counterfactual) rewards. The framework provides policy-level explanations by accessing learned Q-functions and identifying intersecting critical states. Global explanations are generated to summarise policy behaviour through the visualisation of sub-trajectories based on these states, while local explanations are based on the action-values in states. We conduct experiments on several grid-world examples. Our results show that it is possible to explain the difference between learned policies based on Q-functions. This demonstrates the potential for more informed human decision-making when deploying policies and highlights the possibility of developing further XAI techniques in RL.