Exploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis

Exploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Martin Pawelczyk;Chirag Agarwal;Shalmali Joshi;Sohini Upadhyay;Himabindu Lakkaraju
Martin Pawelczyk;Chirag Agarwal;Shalmali Joshi;Sohini Upadhyay;Himabindu Lakkaraju
中科院分区:
其他
文献类型:
--
作者:
Martin Pawelczyk;Chirag Agarwal;Shalmali Joshi;Sohini Upadhyay;Himabindu Lakkaraju

文献摘要

被引文献

相似文献

随着机器学习(ML)模型越来越广泛地部署在高风险的应用程序中,反事实解释已经成为在实践中提供可操作的模型解释的关键工具。尽管反事实解释越来越受欢迎,但对这些解释仍然缺乏更深层次的理解。在这项工作中,我们通过对抗性例子的镜头系统地分析反事实解释。我们通过形式化流行的反事实解释和对抗性例子生成方法之间的相似性来实现这一点,当它们等价时,识别条件。然后,我们推导了反事实解释方法输出的解与对抗性实例生成方法之间的距离的上界,并在几个真实世界的数据集上进行了验证。通过在反事实解释和对抗性例子之间建立这些理论和经验上的相似性,我们的工作提出了关于现有反事实解释算法的设计和开发的基本问题。
As machine learning (ML) models become more widely deployed in high-stakes applications, counterfactual explanations have emerged as key tools for providing actionable model explanations in practice. Despite the growing popularity of counterfactual explanations, a deeper understanding of these explanations is still lacking. In this work, we systematically analyze counterfactual explanations through the lens of adversarial examples. We do so by formalizing the similarities between popular counterfactual explanation and adversarial example generation methods identifying conditions when they are equivalent. We then derive the upper bounds on the distances between the solutions output by counterfactual explanation and adversarial example generation methods, which we validate on several real-world data sets. By establishing these theoretical and empirical similarities between counterfactual explanations and adversarial examples, our work raises fundamental questions about the design and development of existing counterfactual explanation algorithms.