"Is your explanation stable?": A Robustness Evaluation Framework for Feature Attribution

"Is your explanation stable?": A Robustness Evaluation Framework for Feature Attribution
复制标题

DOI:
10.1145/3548606.3559392
复制
发表时间:
2022-09
期刊:
Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Yuyou Gan;Yuhao Mao;Xuhong Zhang;S. Ji;Yuwen Pu;Meng Han;Jianwei Yin;Ting Wang
Yuyou Gan;Yuhao Mao;Xuhong Zhang;S. Ji;Yuwen Pu;Meng Han;Jianwei Yin;Ting Wang
中科院分区:
其他
文献类型:
--
作者:
Yuyou Gan;Yuhao Mao;Xuhong Zhang;S. Ji;Yuwen Pu;Meng Han;Jianwei Yin;Ting Wang

文献摘要

相似文献

神经网络变得越来越流行。然而,了解他们的决策过程却很复杂。解释模型行为的一种重要方法是特征归因,即将其决策归因于关键特征。尽管提出了许多算法,但大多数算法都是为了提高模型的忠实度(保真度)。然而,真实环境中包含许多随机噪声,这可能会导致相似图像的特征归因图受到很大的扰动。更严重的是,最近的研究表明,解释算法很容易受到对抗性攻击,为恶意扰动的输入生成相同的解释。所有这些都使得解释在实际场景中难以相信,特别是在安全关键型应用程序中。为了弥补这一差距,我们提出了特征归因中值测试(MeTFA)来量化不确定性并通过理论保证提高解释算法的稳定性。 MeTFA 与方法无关,即它可以应用于任何特征归因方法。 MeTFA具有以下两个功能:(1)检查一个特征是否显着重要或不重要,并生成MeTFA显着图以可视化结果; (2) 计算特征归因得分的置信区间并生成 MeTFA 平滑图以增加解释的稳定性。大量实验表明,MeTFA 提高了解释的视觉质量,并显着降低了不稳定性,同时保持了原始方法的忠实度。为了定量评估 MeTFA 的可信度和稳定性,我们进一步提出了几个鲁棒的可信度指标,可以评估不同噪声设置下解释的可​​信度。实验结果表明,MeTFA 平滑解释可以显着提高鲁棒可信度。此外,我们还通过两个典型应用来展示MeTFA在应用中的潜力。首先,当应用于 SOTA 解释方法来定位语义分割模型的上下文偏差时,MeTFA 显着解释使用小得多的区域来保持 99% 以上的忠实度。其次,在测试不同的面向解释的攻击时,MeTFA 可以帮助防御普通攻击以及针对解释的自适应对抗性攻击。
Neural networks have become increasingly popular. Nevertheless, understanding their decision process turns out to be complicated. One vital method to explain a models' behavior is feature attribution, i.e., attributing its decision to pivotal features. Although many algorithms are proposed, most of them aim to improve the faithfulness (fidelity) to the model. However, the real environment contains many random noises, which may cause the feature attribution maps to be greatly perturbed for similar images. More seriously, recent works show that explanation algorithms are vulnerable to adversarial attacks, generating the same explanation for a maliciously perturbed input. All of these make the explanation hard to trust in real scenarios, especially in security-critical applications. To bridge this gap, we propose Median Test for Feature Attribution (MeTFA) to quantify the uncertainty and increase the stability of explanation algorithms with theoretical guarantees. MeTFA is method-agnostic, i.e., it can be applied to any feature attribution method. MeTFA has the following two functions: (1) examine whether one feature is significantly important or unimportant and generate a MeTFA-significant map to visualize the results; (2) compute the confidence interval of a feature attribution score and generate a MeTFA-smoothed map to increase the stability of the explanation. Extensive experiments show that MeTFA improves the visual quality of explanations and significantly reduces the instability while maintaining the faithfulness of the original method. To quantitatively evaluate MeTFA's faithfulness and stability, we further propose several robust faithfulness metrics, which can evaluate the faithfulness of an explanation under different noise settings. Experiment results show that the MeTFA-smoothed explanation can significantly increase the robust faithfulness. In addition, we use two typical applications to show MeTFA's potential in the applications. First, when being applied to the SOTA explanation method to locate context bias for semantic segmentation models, MeTFA-significant explanations use far smaller regions to maintain 99%+ faithfulness. Second, when testing with different explanation-oriented attacks, MeTFA can help defend vanilla, as well as adaptive, adversarial attacks against explanations.