BlindNet backdoor: Attack on deep neural network using blind watermark

BlindNet backdoor: Attack on deep neural network using blind watermark
复制标题

DOI:
10.1007/s11042-021-11135-0
复制
发表时间:
2022-01-07
影响因子:
3.6
通讯作者:
Kim, Yongchul
Kim, Yongchul
中科院分区:
计算机科学4区
文献类型:
--
作者:
Kwon, Hyun;Kim, Yongchul

文献摘要

被引文献

相似文献

深度神经网络(DNN)在图像识别、语音识别、视频识别和模式分析方面具有出色的性能。然而,DNN容易受到后门攻击。后门攻击允许DNN正确识别不包含特定触发器的正常数据,但会导致它错误地识别包含触发器的数据。后门攻击的一个优点是,攻击者可以通过使用特定的触发器来确定攻击时间。在本文中,我们提出了一个盲水印后门方法,其结果是不可感知的人类。与现有方法不同,该方法通过使触发器不可见来避免后门样本攻击的人类可检测性。在该方法中,通过使用傅立叶变换将由频带中的特定图像组成的触发器插入到输入数据中来生成盲水印样本。通过在训练过程中对盲水印样本进行额外的训练,目标模型学会了错误地对具有特定水印的任何样本进行分类。为了进行测试,我们使用了CIFAR 10数据集和Tensorflow机器学习库。在实验中,当盲水印样本在训练数据中所占比例为10%时,该模型对原始样本的分类正确率为88.9%,利用盲水印样本进行训练的攻击成功率为99.3%。
Deep neural network (DNNs) provide excellent performance in image recognition, speech recognition, video recognition, and pattern analysis. However, DNNs are vulnerable to backdoor attacks. A backdoor attack allows a DNN to correctly recognize normal data that do not contain a specific trigger but induces it to incorrectly recognize data that do contain the trigger. An advantage of the backdoor attack is that the attacker can determine the time of attack by using a specific trigger. In this paper, we propose a blind-watermark backdoor method whose results are imperceptible to humans. Unlike existing methods, the proposed method avoids the human detectability of the backdoor sample attack by making the trigger invisible. In this method, a blind-watermarked sample is generated by inserting a trigger consisting of a specific image in a frequency band into input data by using a Fourier transform. By additionally training on the blind-watermarked sample during the training process, the target model learns to incorrectly classify any sample with the specific watermark. For testing, we used the CIFAR10 dataset and the Tensorflow machine learning library. In the experiment, when the proportion of blind-watermarked samples in the training data was 10%, the proposed method resulted in 88.9% classification accuracy by the model on the original samples and a 99.3% attack success rate via training with the blind-watermarked samples.