Attribution-based Explanations that Provide Recourse Cannot be Robust

Attribution-based Explanations that Provide Recourse Cannot be Robust
复制标题

DOI:
10.48550/arxiv.2205.15834
复制
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
通讯作者:
H. Fokkema;R. D. Heide;T. Erven
H. Fokkema;R. D. Heide;T. Erven
中科院分区:
其他
文献类型:
--
作者:
H. Fokkema;R. D. Heide;T. Erven

文献摘要

被引文献

相似文献

机器学习方法的不同用户需要不同的解释,这取决于他们的目标。为了使机器学习对社会负责,一个重要的目标是获得可操作的追索权选项,这允许受影响的用户通过对其输入$x$进行有限的更改来更改机器学习系统的决策$f(x)$。我们通过提供追索权敏感性的一般定义来形式化这一点,该定义需要用效用函数来实例化,该效用函数描述决策的哪些变化与用户相关。该定义适用于局部属性方法,该方法将重要性权重赋予每个输入特征。通常认为,这种局部属性应该是鲁棒的,在这个意义上,被解释的输入$x$的小变化不应该导致特征权重的大变化。然而,我们正式证明,它是在一般情况下不可能的任何单一的归属方法,既追索权敏感和强大的同时。因此,这些性质中至少有一个必须总是存在反例。我们为几种流行的归因方法提供了这样的反例,包括LIME,SHAP,IntegratedConcentrants和SmoothGrad。我们的研究结果还涵盖了反事实的解释,这可能被视为归因,描述了一个扰动的x$。我们进一步讨论了解决我们的不可能结果的可能方法,例如允许输出由具有多个属性的集合组成,并且我们为特定类别的连续函数提供了充分条件,使其对追索权敏感。最后,我们加强了我们的不可能性结果的限制情况下,用户只能改变一个属性的$x$,通过提供一个精确的表征的功能$f$的不可能性适用。
Different users of machine learning methods require different explanations, depending on their goals. To make machine learning accountable to society, one important goal is to get actionable options for recourse, which allow an affected user to change the decision $f(x)$ of a machine learning system by making limited changes to its input $x$. We formalize this by providing a general definition of recourse sensitivity, which needs to be instantiated with a utility function that describes which changes to the decisions are relevant to the user. This definition applies to local attribution methods, which attribute an importance weight to each input feature. It is often argued that such local attributions should be robust, in the sense that a small change in the input $x$ that is being explained, should not cause a large change in the feature weights. However, we prove formally that it is in general impossible for any single attribution method to be both recourse sensitive and robust at the same time. It follows that there must always exist counterexamples to at least one of these properties. We provide such counterexamples for several popular attribution methods, including LIME, SHAP, Integrated Gradients and SmoothGrad. Our results also cover counterfactual explanations, which may be viewed as attributions that describe a perturbation of $x$. We further discuss possible ways to work around our impossibility result, for instance by allowing the output to consist of sets with multiple attributions, and we provide sufficient conditions for specific classes of continuous functions to be recourse sensitive. Finally, we strengthen our impossibility result for the restricted case where users are only able to change a single attribute of $x$, by providing an exact characterization of the functions $f$ to which impossibility applies.