ADDMU: Detection of Far-Boundary Adversarial Examples with Data and Model Uncertainty Estimation

ADDMU: Detection of Far-Boundary Adversarial Examples with Data and Model Uncertainty Estimation
复制标题

DOI:
10.48550/arxiv.2210.12396
复制
发表时间:
2022-10
期刊:
--
影响因子:
--
通讯作者:
Fan Yin;Yao Li;Cho-Jui Hsieh;Kai-Wei Chang
Fan Yin;Yao Li;Cho-Jui Hsieh;Kai-Wei Chang
中科院分区:
其他
文献类型:
--
作者:
Fan Yin;Yao Li;Cho-Jui Hsieh;Kai-Wei Chang

文献摘要

相似文献

对抗样例检测(AED)是对抗对抗攻击的关键防御技术,近年来受到自然语言处理(NLP)界的广泛关注。尽管新的AED方法层出不穷,但我们的研究表明,现有的方法严重依赖于捷径来获得良好的效果。换句话说,目前NLP中基于搜索的对抗性攻击一旦模型预测发生变化就会停止,因此这些攻击产生的大多数对抗性示例都位于模型决策边界附近。为了超越这种捷径并公平地评估AED方法,我们提出用远边界(FB)对抗示例来测试AED方法。在这种情况下,现有的方法表现出比随机猜测更差的性能。为了克服这一限制,我们提出了一种新的技术,ADDMU,具有数据和模型不确定性的对手检测,它结合了常规和FB对抗样本检测的两种不确定性估计。在每种情况下,我们的新方法比以前的方法分别高出3.6和6.0 AUC点。最后,我们的分析表明,ADDMU提供的两种类型的不确定性可以用来表征对抗性示例,并识别对抗性训练中对模型鲁棒性贡献最大的示例。
Adversarial Examples Detection (AED) is a crucial defense technique against adversarial attacks and has drawn increasing attention from the Natural Language Processing (NLP) community. Despite the surge of new AED methods, our studies show that existing methods heavily rely on a shortcut to achieve good performance. In other words, current search-based adversarial attacks in NLP stop once model predictions change, and thus most adversarial examples generated by those attacks are located near model decision boundaries. To surpass this shortcut and fairly evaluate AED methods, we propose to test AED methods with Far Boundary (FB) adversarial examples. Existing methods show worse than random guess performance under this scenario. To overcome this limitation, we propose a new technique, ADDMU, adversary detection with data and model uncertainty, which combines two types of uncertainty estimation for both regular and FB adversarial example detection. Our new method outperforms previous methods by 3.6 and 6.0 AUC points under each scenario. Finally, our analysis shows that the two types of uncertainty provided by ADDMU can be leveraged to characterize adversarialexamples and identify the ones that contribute most to model’s robustness in adversarial training.