CX-ToM: Counterfactual explanations with theory-of-mind for enhancing human trust in image recognition models.

CX-ToM: Counterfactual explanations with theory-of-mind for enhancing human trust in image recognition models.
复制标题

DOI:
10.1016/j.isci.2021.103581
复制
发表时间:
2022-01-21
期刊:
影响因子:
5.8
通讯作者:
Zhu SC
Zhu SC
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Akula AR;Wang K;Liu C;Saba-Sadiya S;Lu H;Todorovic S;Chai J;Zhu SC

文献摘要

参考文献

被引文献

相似文献

我们提出了CX-ToM,即反事实解释的缩写,这是一种新的可解释人工智能(XAI)框架,用于解释由深度卷积神经网络(CNN)做出的决策。与XAI中当前将解释作为单次响应生成的方法相反,我们将解释作为一个迭代的通信过程,即机器和人类用户之间的对话。更具体地说,我们的CX-ToM框架通过调解机器和人类用户思想之间的差异,在对话中生成一系列解释。为了做到这一点,我们使用心智理论(ToM),它帮助我们明确地建模人类的意图,由人类推断的机器思想,以及由机器推断的人类思想。此外,大多数最先进的XAI框架都提供了基于注意力(或热图)的解释。在我们的工作中,我们表明这些基于注意力的解释不足以增加人类对底层CNN模型的信任。在CX-ToM中,我们使用了称为断层线的反事实解释,我们将其定义如下:给定一个输入图像I, CNN分类模型M预测类别cpred,断层线识别最小语义级特征(例如,斑马上的条纹),称为可解释的概念,需要从I中添加或删除,以将M的分类类别I更改为另一个指定的类别cald。大量的实验验证了我们的假设,表明我们的CX-ToM显著优于最先进的XAI模型。注意不是一个好的解释解释是一个互动的沟通过程我们介绍了一个新的基于心理理论和反事实解释的XAI框架。计算机科学;人工智能;人机交互
We propose CX-ToM, short for counterfactual explanations with theory-of-mind, a new explainable AI (XAI) framework for explaining decisions made by a deep convolutional neural network (CNN). In contrast to the current methods in XAI that generate explanations as a single shot response, we pose explanation as an iterative communication process, i.e., dialogue between the machine and human user. More concretely, our CX-ToM framework generates a sequence of explanations in a dialogue by mediating the differences between the minds of the machine and human user. To do this, we use Theory of Mind (ToM) which helps us in explicitly modeling the human’s intention, the machine’s mind as inferred by the human, as well as human's mind as inferred by the machine. Moreover, most state-of-the-art XAI frameworks provide attention (or heat map) based explanations. In our work, we show that these attention-based explanations are not sufficient for increasing human trust in the underlying CNN model. In CX-ToM, we instead use counterfactual explanations called fault-lines which we define as follows: given an input image I for which a CNN classification model M predicts class cpred, a fault-line identifies the minimal semantic-level features (e.g., stripes on zebra), referred to as explainable concepts, that need to be added to or deleted from I to alter the classification category of I by M to another specified class calt. Extensive experiments verify our hypotheses, demonstrating that our CX-ToM significantly outperforms the state-of-the-art XAI models. Attention is not a Good Explanation Explanation is an Interactive Communication Process We introduce a new XAI framework based on Theory-of-Mind and counterfactual explana- tions. Computer science; Artificial intelligence; Human-computer interaction
DOI: 10.1080/14640748708401804
发表时间: 1987-11-01
期刊: QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY SECTION A-HUMAN EXPERIMENTAL PSYCHOLOGY
影响因子: --
作者:
BERRY, DC;BROADBENT, DE
通讯作者: BROADBENT, DE
DOI: 10.1007/s11063-011-9207-8
发表时间: 2012-04-01
影响因子: 3.1
作者:
Augasta, M. Gethsiyal;Kathirvalavakumar, T.
通讯作者: Kathirvalavakumar, T.
DOI: 10.1137/080716542
发表时间: 2009-01-01
影响因子: 2.1
作者:
Beck, Amir;Teboulle, Marc
通讯作者: Teboulle, Marc
DOI: 10.1147/jrd.2016.2629318
发表时间: 2017-01-01
影响因子: 1.3
作者:
Agarwal, S.;Aggarwal, V.;Sridhara, G.
通讯作者: Sridhara, G.
DOI: 10.1371/journal.pone.0130140
发表时间: 2015
期刊: PloS one
影响因子: 3.7
作者:
Bach S;Binder A;Montavon G;Klauschen F;Müller KR;Samek W
通讯作者: Samek W