Towards Automating Model Explanations with Certified Robustness Guarantees

Towards Automating Model Explanations with Certified Robustness Guarantees
复制标题

DOI:
10.1609/aaai.v36i6.20651
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Mengdi Huai;Jinduo Liu;Chenglin Miao;Liuyi Yao;Aidong Zhang
Mengdi Huai;Jinduo Liu;Chenglin Miao;Liuyi Yao;Aidong Zhang
中科院分区:
其他
文献类型:
--
作者:
Mengdi Huai;Jinduo Liu;Chenglin Miao;Liuyi Yao;Aidong Zhang

文献摘要

相似文献

最近,提供模型解释变得非常流行。与传统的特征层次模型解释相比,基于概念的解释可以提供高级人类概念形式的解释。然而,现有的基于概念的解释方法隐含地遵循两步程序,涉及人类干预。具体来说,他们首先需要人类参与定义(或提取)高级概念,然后以事后方式手动计算这些识别概念的重要性分数。由于手工工作,这一费力的过程需要大量的人力和资源支出,这阻碍了它们的大规模部署。在实践中,由于定义基于概念的可解释性单元的主观性,在没有人工干预的情况下自动生成基于概念的解释是具有挑战性的。此外,由于其数据驱动的性质,可解释性本身也可能容易受到恶意操纵。因此,本文的目标是将人类从这个繁琐的过程中解放出来,同时确保生成的解释对对抗性扰动具有可证明的鲁棒性。我们提出了一种新的基于概念的解释方法,它不仅可以自动提供基于原型的概念解释,但也提供认证的鲁棒性保证生成的基于原型的解释。我们还进行了广泛的实验,在现实世界的数据集,以验证所提出的方法的理想属性。
Providing model explanations has gained significant popularity recently. In contrast with the traditional feature-level model explanations, concept-based explanations can provide explanations in the form of high-level human concepts. However, existing concept-based explanation methods implicitly follow a two-step procedure that involves human intervention. Specifically, they first need the human to be involved to define (or extract) the high-level concepts, and then manually compute the importance scores of these identified concepts in a post-hoc way. This laborious process requires significant human effort and resource expenditure due to manual work, which hinders their large-scale deployability. In practice, it is challenging to automatically generate the concept-based explanations without human intervention due to the subjectivity of defining the units of concept-based interpretability. In addition, due to its data-driven nature, the interpretability itself is also potentially susceptible to malicious manipulations. Hence, our goal in this paper is to free human from this tedious process, while ensuring that the generated explanations are provably robust to adversarial perturbations. We propose a novel concept-based interpretation method, which can not only automatically provide the prototype-based concept explanations but also provide certified robustness guarantees for the generated prototype-based explanations. We also conduct extensive experiments on real-world datasets to verify the desirable properties of the proposed method.