A Synergetic Attack against Neural Network Classifiers combining Backdoor and Adversarial Examples

A Synergetic Attack against Neural Network Classifiers combining Backdoor and Adversarial Examples
复制标题

DOI:
10.1109/bigdata52589.2021.9671964
复制
发表时间:
2021-09
期刊:
2021 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Guanxiong Liu;Issa M. Khalil;Abdallah Khreishah;Nhathai Phan
Guanxiong Liu;Issa M. Khalil;Abdallah Khreishah;Nhathai Phan
中科院分区:
其他
文献类型:
--
作者:
Guanxiong Liu;Issa M. Khalil;Abdallah Khreishah;Nhathai Phan

文献摘要

被引文献

相似文献

神经网络(NN)在关键计算机视觉和图像处理应用中的普遍性使其对对抗性操作非常有吸引力。现有的大量研究彻底调查了两大类针对NN模型完整性的攻击。第一类攻击,通常称为对抗性示例,通过仔细地向输入示例中添加噪声来扰乱模型的推理。在第二类攻击中,攻击者试图在训练过程中通过植入特洛伊木马后门来操纵模型。研究人员表明,这种攻击对日益增长的NN应用构成了严重威胁,并分别针对每种攻击类型提出了几种防御措施。然而,这种片面的防御方法留下潜在的未知风险,在现实世界中的情况下,当对手可以统一不同的攻击,以创建新的和更致命的绕过现有的defenses.In这项工作中,我们展示了如何联合利用对抗扰动和模型中毒漏洞,实际上发动一个新的隐形攻击,称为AdvTrojan。AdvTrojan是隐形的,因为它只能在以下情况下被激活:1)在推理过程中将精心制作的对抗扰动注入输入示例,2)在模型的训练过程中植入特洛伊木马后门。我们利用输入空间中的对抗性噪声将木马感染的示例移动到模型决策边界,使其难以检测。AdvTrojan的隐形行为欺骗用户意外地信任受感染的模型作为对抗对抗性示例的鲁棒分类器。AdvTrojan可以通过类似于传统特洛伊木马后门攻击的训练数据中毒来实现。我们对几个基准数据集的全面分析和广泛实验表明,AdvTrojan可以绕过现有的防御,在我们的大多数实验场景中成功率接近100%,并且可以扩展到攻击联邦学习和高分辨率图像。
The pervasiveness of neural networks (NNs) in critical computer vision and image processing applications makes them very attractive for adversarial manipulation. A large body of existing research thoroughly investigates two broad categories of attacks targeting the integrity of NN models. The first category of attacks, commonly called Adversarial Examples, perturbs the model’s inference by carefully adding noise into input examples. In the second category of attacks, adversaries try to manipulate the model during the training process by implanting Trojan backdoors. Researchers show that such attacks pose severe threats to the growing applications of NNs and propose several defenses against each attack type individually. However, such one-sided defense approaches leave potentially unknown risks in real-world scenarios when an adversary can unify different attacks to create new and more lethal ones bypassing existing defenses.In this work, we show how to jointly exploit adversarial perturbation and model poisoning vulnerabilities to practically launch a new stealthy attack, dubbed AdvTrojan. AdvTrojan is stealthy because it can be activated only when: 1) a carefully crafted adversarial perturbation is injected into the input examples during inference, and 2) a Trojan backdoor is implanted during the training process of the model. We leverage adversarial noise in the input space to move Trojan-infected examples across the model decision boundary, making it difficult to detect. The stealthiness behavior of AdvTrojan fools the users into accidentally trusting the infected model as a robust classifier against adversarial examples. AdvTrojan can be implemented by only poisoning the training data similar to conventional Trojan backdoor attacks. Our thorough analysis and extensive experiments on several benchmark datasets show that AdvTrojan can bypass existing defenses with a success rate close to 100% in most of our experimental scenarios and can be extended to attack federated learning as well as high-resolution images.