Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End

Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End
复制标题

DOI:
10.1145/3461702.3462597
复制
发表时间:
2020-11
期刊:
Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society
影响因子:
--
通讯作者:
R. Mothilal;Divyat Mahajan;Chenhao Tan;Amit Sharma
R. Mothilal;Divyat Mahajan;Chenhao Tan;Amit Sharma
中科院分区:
其他
文献类型:
--
作者:
R. Mothilal;Divyat Mahajan;Chenhao Tan;Amit Sharma

文献摘要

被引文献

相似文献

特征归因和反事实解释是解释ML模型的流行方法。前者为每个输入特征分配一个重要性分数,而后者提供具有最小变化的输入示例,以改变模型的预测。为了统一这些方法,我们提供了一个解释的基础上,实际的因果关系框架,并提出了两个关键的结果,在他们的使用。首先,我们提出了一种方法来生成一组反事实的例子的特征归因解释。这些特征属性传达了特征对于改变模型的分类结果的重要性,特别是关于特征的子集对于该改变是否是必要的和/或足够的,而基于属性的方法无法提供。其次,我们展示了如何反事实的例子,可以用来评估其必要性和充分性的归因为基础的解释的好处。因此,我们强调这两种方法的互补性。我们对三个基准数据集- Adult-Income,LendingClub和German-Credit -的评估证实了这一点。像LIME和SHAP这样的特征归因方法以及像Wachter等人和DiCE这样的反事实解释方法通常在特征重要性排名上不一致。此外,通过限制可以修改以生成反事实示例的特征,我们发现LIME或SHAP的前k个特征通常既不是模型预测的必要解释,也不是充分解释。最后,我们提出了一个案例研究不同的解释方法对现实世界的医院分诊问题。
Feature attributions and counterfactual explanations are popular approaches to explain a ML model. The former assigns an importance score to each input feature, while the latter provides input examples with minimal changes to alter the model's predictions. To unify these approaches, we provide an interpretation based on the actual causality framework and present two key results in terms of their use. First, we present a method to generate feature attribution explanations from a set of counterfactual examples. These feature attributions convey how important a feature is to changing the classification outcome of a model, especially on whether a subset of features is necessary and/or sufficient for that change, which attribution-based methods are unable to provide. Second, we show how counterfactual examples can be used to evaluate the goodness of an attribution-based explanation in terms of its necessity and sufficiency. As a result, we highlight the complimentary of these two approaches. Our evaluation on three benchmark datasets --- Adult-Income, LendingClub, and German-Credit --- confirms the complimentary. Feature attribution methods like LIME and SHAP and counterfactual explanation methods like Wachter et al. and DiCE often do not agree on feature importance rankings. In addition, by restricting the features that can be modified for generating counterfactual examples, we find that the top-k features from LIME or SHAP are often neither necessary nor sufficient explanations of a model's prediction. Finally, we present a case study of different explanation methods on a real-world hospital triage problem.