CLEAR: Clean-up Sample-Targeted Backdoor in Neural Networks

CLEAR: Clean-up Sample-Targeted Backdoor in Neural Networks
复制标题

DOI:
10.1109/iccv48922.2021.01614
复制
发表时间:
2021-10
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Liuwan Zhu;R. Ning;Chunsheng Xin;Chong Wang;Hongyi Wu
Liuwan Zhu;R. Ning;Chunsheng Xin;Chong Wang;Hongyi Wu
中科院分区:
其他
文献类型:
--
作者:
Liuwan Zhu;R. Ning;Chunsheng Xin;Chong Wang;Hongyi Wu

文献摘要

相似文献

数据中毒攻击引发了对深度神经网络安全的严重安全担忧,因为它可能导致神经后门,错误分类攻击者精心设计的某些输入。特别是,针对样本的后门攻击是一个新的挑战。它针对一个或几个特定的样本,称为目标样本,以将它们错误地分类为目标类别。如果没有在后门模型中植入触发器,现有的后门检测方案无法检测到样本目标后门,因为它们依赖于对触发器或触发器的强大功能进行逆向工程。在本文中,我们提出了一种新的方案来检测和缓解样本目标后门攻击。我们发现并演示了样本目标后门的一个独特属性,它迫使边界改变,从而在目标样本周围形成小的“口袋”。基于这一观察结果,我们提出了一种新的防御机制,通过将恶意口袋“包装”到特征空间中的一个紧致凸壳中来精确定位恶意口袋。我们设计了一种有效的算法来搜索这样的凸壳,并根据凸壳对识别出的带有正确标签的恶意样本进行模型微调,从而去除后门。实验表明,该方法对于检测和缓解大范围的样本目标后门攻击具有很高的效率。
The data poisoning attack has raised serious security concerns on the safety of deep neural networks, since it can lead to neural backdoor that misclassifies certain inputs crafted by an attacker. In particular, the sample-targeted backdoor attack is a new challenge. It targets at one or a few specific samples, called target samples, to misclassify them to a target class. Without a trigger planted in the backdoor model, the existing backdoor detection schemes fail to detect the sample-targeted backdoor as they depend on reverse-engineering the trigger or strong features of the trigger. In this paper, we propose a novel scheme to detect and mitigate sample-targeted backdoor attacks. We discover and demonstrate a unique property of the sample-targeted backdoor, which forces a boundary change such that small "pockets" are formed around the target sample. Based on this observation, we propose a novel defense mechanism to pinpoint a malicious pocket by "wrapping" them into a tight convex hull in the feature space. We design an effective algorithm to search for such a convex hull and remove the backdoor by fine-tuning the model using the identified malicious samples with the corrected label according to the convex hull. The experiments show that the proposed approach is highly efficient for detecting and mitigating a wide range of sample-targeted backdoor attacks.