"Why Not Other Classes?": Towards Class-Contrastive Back-Propagation Explanations

"Why Not Other Classes?": Towards Class-Contrastive Back-Propagation Explanations
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Yipei Wang;Xiaoqian Wang
Yipei Wang;Xiaoqian Wang
中科院分区:
其他
文献类型:
--
作者:
Yipei Wang;Xiaoqian Wang

文献摘要

相似文献

已经开发了许多方法来解释基于深度神经网络(DNN)的分类器的内部机制。现有的解释方法往往局限于解释一个预先指定的类的预测,这回答了“为什么输入被归类到这个类?”然而,这样的解释相对于一个单一的类本质上是不够的,因为他们不捕捉功能与类的区分能力。也就是说,对于预测一个类别重要的特征也可能对于其他类别重要。为了捕获具有真正类别区分能力的特征,我们应该问“为什么输入被分类到这个类别,而不是其他类别?”为了回答这个问题,我们提出了一个解释DNN的加权对比框架。我们的框架可以很容易地转换任何现有的反向传播解释方法,以建立类对比解释。我们从理论上验证了我们的加权对比解释一般的反向传播解释,并表明,我们的框架,使类对比的解释与定性和定量实验的显着改善。基于这些结果,我们指出了当前可解释人工智能(XAI)研究中的一个重要盲点,即对预测逻辑和概率的解释是模糊的。我们建议在任何解释方法的应用中,都应明确区分这两个方面。
Numerous methods have been developed to explain the inner mechanism of deep neural network (DNN) based classifiers. Existing explanation methods are often limited to explaining predictions of a pre-specified class, which answers the question “why is the input classified into this class?” However, such explanations with respect to a single class are inherently insufficient because they do not capture features with class-discriminative power. That is, features that are important for predicting one class may also be important for other classes. To capture features with true class-discriminative power, we should instead ask “why is the input classified into this class, but not others ?” To answer this question, we propose a weighted contrastive framework for explaining DNNs. Our framework can easily convert any existing back-propagation explanation methods to build class-contrastive explanations. We theoretically validate our weighted contrast explanation in general back-propagation explanations, and show that our framework enables class-contrastive explanations with significant improvements in both qualitative and quantitative experiments. Based on the results, we point out an important blind spot in the current explainable artificial intelligence (XAI) study, where explanations towards the predicted logits and the probabilities are obfus-cated. We suggest that these two aspects should be distinguished explicitly any time explanation methods are applied.