Fairwashing Explanations with Off-Manifold Detergent

Fairwashing Explanations with Off-Manifold Detergent
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Christopher J. Anders;Plamen Pasliev;Ann-Kathrin Dombrowski;K. Müller;P. Kessel
Christopher J. Anders;Plamen Pasliev;Ann-Kathrin Dombrowski;K. Müller;P. Kessel
中科院分区:
其他
文献类型:
--
作者:
Christopher J. Anders;Plamen Pasliev;Ann-Kathrin Dombrowski;K. Müller;P. Kessel

文献摘要

被引文献

相似文献

解释方法有望使黑盒分类器更加透明。因此,希望它们可以作为算法的合理,公平和值得信赖的决策过程的证明,从而提高最终用户的接受度。在本文中,我们从理论上和实验上表明,这些希望目前是没有根据的。具体来说,我们表明,对于任何分类器g,我们总是可以构建另一个分类器g,它对数据具有相同的行为(相同的训练,验证和测试错误),但任意操纵解释图。我们使用微分几何理论推导出这一声明,并通过实验证明了各种解释方法,架构和数据集。我们的理论见解的动机,然后,我们提出了一个修改现有的解释方法,使他们显着更强大。
Explanation methods promise to make black-box classifiers more transparent. As a result, it is hoped that they can act as proof for a sensible, fair and trustworthy decision-making process of the algorithm and thereby increase its acceptance by the end-users. In this paper, we show both theoretically and experimentally that these hopes are presently unfounded. Specifically, we show that, for any classifier $g$, one can always construct another classifier $\tilde{g}$ which has the same behavior on the data (same train, validation, and test error) but has arbitrarily manipulated explanation maps. We derive this statement theoretically using differential geometry and demonstrate it experimentally for various explanation methods, architectures, and datasets. Motivated by our theoretical insights, we then propose a modification of existing explanation methods which makes them significantly more robust.