Explaining by Removing: A Unified Framework for Model Explanation

Explaining by Removing: A Unified Framework for Model Explanation
复制标题

DOI:
--
复制
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Ian Covert;Scott M. Lundberg;Su-In Lee
Ian Covert;Scott M. Lundberg;Su-In Lee
中科院分区:
其他
文献类型:
--
作者:
Ian Covert;Scott M. Lundberg;Su-In Lee

文献摘要

被引文献

相似文献

研究人员已经提出了各种各样的模型解释方法,但仍然不清楚大多数方法是如何相关的,或者一种方法何时优于另一种方法。我们建立了一类新的方法,去除为基础的解释,是基于模拟功能去除的原则,以量化每个功能的影响。这些方法在几个方面有所不同,因此我们开发了一个框架,该框架沿沿着三个维度表征每个方法:1)该方法如何移除特征,2)该方法解释什么模型行为,以及3)该方法如何总结每个特征的影响。我们的框架统一了25种现有的方法,包括几种最广泛使用的方法(SHAP,LIME,有意义的扰动,排列测试)。这类新的解释方法有着丰富的联系,我们使用的工具在很大程度上被可解释性文献所忽视。为了在认知心理学中基于锚移除的解释,我们证明了特征移除是减法反事实推理的一个简单应用。合作博弈论的思想揭示了不同方法之间的关系和权衡,我们推导出的条件下,所有基于删除的解释有信息理论的解释。通过这种分析,我们开发了一个统一的框架,帮助从业者更好地理解模型解释工具,并提供了一个强大的理论基础,未来的可解释性研究可以建立。
Researchers have proposed a wide variety of model explanation approaches, but it remains unclear how most methods are related or when one method is preferable to another. We establish a new class of methods, removal-based explanations, that are based on the principle of simulating feature removal to quantify each feature's influence. These methods vary in several respects, so we develop a framework that characterizes each method along three dimensions: 1) how the method removes features, 2) what model behavior the method explains, and 3) how the method summarizes each feature's influence. Our framework unifies 25 existing methods, including several of the most widely used approaches (SHAP, LIME, Meaningful Perturbations, permutation tests). This new class of explanation methods has rich connections that we examine using tools that have been largely overlooked by the explainability literature. To anchor removal-based explanations in cognitive psychology, we show that feature removal is a simple application of subtractive counterfactual reasoning. Ideas from cooperative game theory shed light on the relationships and trade-offs among different methods, and we derive conditions under which all removal-based explanations have information-theoretic interpretations. Through this analysis, we develop a unified framework that helps practitioners better understand model explanation tools, and that offers a strong theoretical foundation upon which future explainability research can build.