Hidden Trigger Backdoor Attacks

Hidden Trigger Backdoor Attacks
复制标题

DOI:
10.1609/aaai.v34i07.6871
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Aniruddha Saha;Akshayvarun Subramanya;H. Pirsiavash
Aniruddha Saha;Akshayvarun Subramanya;H. Pirsiavash
中科院分区:
其他
文献类型:
--
作者:
Aniruddha Saha;Akshayvarun Subramanya;H. Pirsiavash

文献摘要

被引文献

相似文献

随着深度学习算法在各个领域的成功,研究对抗性攻击以保护真实的世界应用中的深度模型已经成为一个重要的研究课题。后门攻击是对深度网络的一种对抗性攻击,攻击者向受害者提供有毒数据来训练模型,然后通过在测试时显示特定的小触发模式来激活攻击。大多数最先进的后门攻击要么提供可以通过视觉检查识别的错误标记的中毒数据,揭示中毒数据中的触发器,要么使用噪声来隐藏触发器。我们提出了一种新形式的后门攻击,其中中毒数据看起来很自然,具有正确的标签,更重要的是,攻击者将触发器隐藏在中毒数据中,并将触发器保密,直到测试时间。我们对各种图像分类设置进行了广泛的研究,并表明我们的攻击可以通过将触发器粘贴在不可见图像上的随机位置来欺骗模型,尽管该模型在干净数据上表现良好。我们还表明,我们提出的攻击不能很容易地捍卫使用最先进的防御算法后门攻击。
With the success of deep learning algorithms in various domains, studying adversarial attacks to secure deep models in real world applications has become an important research topic. Backdoor attacks are a form of adversarial attacks on deep networks where the attacker provides poisoned data to the victim to train the model with, and then activates the attack by showing a specific small trigger pattern at the test time. Most state-of-the-art backdoor attacks either provide mislabeled poisoning data that is possible to identify by visual inspection, reveal the trigger in the poisoned data, or use noise to hide the trigger. We propose a novel form of backdoor attack where poisoned data look natural with correct labels and also more importantly, the attacker hides the trigger in the poisoned data and keeps the trigger secret until the test time. We perform an extensive study on various image classification settings and show that our attack can fool the model by pasting the trigger at random locations on unseen images although the model performs well on clean data. We also show that our proposed attack cannot be easily defended using a state-of-the-art defense algorithm for backdoor attacks.