Latent Backdoor Attacks on Deep Neural Networks

Latent Backdoor Attacks on Deep Neural Networks
复制标题

DOI:
10.1145/3319535.3354209
复制
发表时间:
2019-11
期刊:
Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Yuanshun Yao;Huiying Li;Haitao Zheng;Ben Y. Zhao
Yuanshun Yao;Huiying Li;Haitao Zheng;Ben Y. Zhao
中科院分区:
其他
文献类型:
--
作者:
Yuanshun Yao;Huiying Li;Haitao Zheng;Ben Y. Zhao

文献摘要

被引文献

相似文献

最近的研究提出了对深度神经网络(dnn)进行后门攻击的概念,其中错误分类规则隐藏在正常模型中,只有在非常特定的输入下才会触发。然而,这些“传统的”后门假定用户从头开始训练他们自己的模型,这在实践中很少发生。相反,用户通常会定制“教师”模型,这些模型已经由谷歌等提供商通过一种称为迁移学习的过程进行了预训练。这种定制过程引入了对模型的重大更改,并破坏了隐藏的后门,大大减少了后门在实践中的实际影响。在本文中,我们描述了潜伏后门,这是一种更强大和隐蔽的后门攻击变体,它在迁移学习下起作用。潜在后门是嵌入在“教师”模型中的不完整后门,并通过迁移学习被多个“学生”模型自动继承。如果任何Student模型包含后门所针对的标签,则其定制过程完成后门并使其处于活动状态。我们证明了潜在后门在各种应用环境中都是非常有效的,并通过对交通标志识别、志愿者虹膜识别和公众人物(政治家)面部识别的真实攻击来验证其实用性。最后,我们评估了4种潜在的防御措施,发现只有一种可以有效地破坏潜在的后门,但可能会在分类精度上产生代价。
Recent work proposed the concept of backdoor attacks on deep neural networks (DNNs), where misclassification rules are hidden inside normal models, only to be triggered by very specific inputs. However, these "traditional" backdoors assume a context where users train their own models from scratch, which rarely occurs in practice. Instead, users typically customize "Teacher" models already pretrained by providers like Google, through a process called transfer learning. This customization process introduces significant changes to models and disrupts hidden backdoors, greatly reducing the actual impact of backdoors in practice. In this paper, we describe latent backdoors, a more powerful and stealthy variant of backdoor attacks that functions under transfer learning. Latent backdoors are incomplete backdoors embedded into a "Teacher" model, and automatically inherited by multiple "Student" models through transfer learning. If any Student models include the label targeted by the backdoor, then its customization process completes the backdoor and makes it active. We show that latent backdoors can be quite effective in a variety of application contexts, and validate its practicality through real-world attacks against traffic sign recognition, iris identification of volunteers, and facial recognition of public figures (politicians). Finally, we evaluate 4 potential defenses, and find that only one is effective in disrupting latent backdoors, but might incur a cost in classification accuracy as tradeoff.