Consistent Counterfactuals for Deep Models

Consistent Counterfactuals for Deep Models
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
E. Black;Zifan Wang;Matt Fredrikson;Anupam Datta
E. Black;Zifan Wang;Matt Fredrikson;Anupam Datta
中科院分区:
其他
文献类型:
--
作者:
E. Black;Zifan Wang;Matt Fredrikson;Anupam Datta

文献摘要

被引文献

相似文献

反事实示例是解释机器学习模型在金融和医疗诊断等关键领域的预测的最常用方法之一。反事实通常是在假设使用它们的模型是静态的情况下讨论的,但是在部署模型中可能会定期重新训练或微调。本文研究了在初始训练条件发生微小变化的情况下,深度网络中反事实示例的模型预测的一致性,例如在模型部署过程中经常发生的权重初始化和数据中的留一变化。我们通过实验证明,深度模型的反事实示例在这种小的变化中通常是不一致的,并且增加反事实的成本,这是先前在简单模型的背景下提出的一种增强稳定性的缓解措施,在深度网络中并不是一种可靠的启发式方法。相反,我们的分析表明,模型在反事实周围的局部Lipschitz连续性是其在相关模型中一致性的关键。为此,我们提出了稳定邻居搜索作为一种生成更一致的反事实解释的方法,并在几个基准数据集上说明了这种方法的有效性。
Counterfactual examples are one of the most commonly-cited methods for explaining the predictions of machine learning models in key areas such as finance and medical diagnosis. Counterfactuals are often discussed under the assumption that the model on which they will be used is static, but in deployment models may be periodically retrained or fine-tuned. This paper studies the consistency of model prediction on counterfactual examples in deep networks under small changes to initial training conditions, such as weight initialization and leave-one-out variations in data, as often occurs during model deployment. We demonstrate experimentally that counterfactual examples for deep models are often inconsistent across such small changes, and that increasing the cost of the counterfactual, a stability-enhancing mitigation suggested by prior work in the context of simpler models, is not a reliable heuristic in deep networks. Rather, our analysis shows that a model's local Lipschitz continuity around the counterfactual is key to its consistency across related models. To this end, we propose Stable Neighbor Search as a way to generate more consistent counterfactual explanations, and illustrate the effectiveness of this approach on several benchmark datasets.