BlurNet: Defense by Filtering the Feature Maps

BlurNet: Defense by Filtering the Feature Maps
复制标题

DOI:
10.1109/dsn-w50199.2020.00016
复制
发表时间:
2019-08
期刊:
2020 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W)
影响因子:
--
通讯作者:
Ravi Raju;Mikko H. Lipasti
Ravi Raju;Mikko H. Lipasti
中科院分区:
其他
文献类型:
--
作者:
Ravi Raju;Mikko H. Lipasti

文献摘要

相似文献

最近,对抗性机器学习领域已经引起了人们的关注,因为它表明最先进的深度神经网络容易受到对抗性示例的影响,这是由于在输入图像中添加了小的扰动。对抗性示例由恶意对手通过获得对模型参数(例如梯度信息)的访问以改变输入或通过攻击替代模型并将这些恶意示例转移到攻击受害者模型来生成。具体而言,这些攻击算法之一,鲁棒物理扰动$(RP_{2})$,生成黑色和白色贴纸的停止标志的对抗图像,以实现针对标准架构的交通标志分类器的高目标误分类率。在本文中,我们提出了BlurNet,一种针对RP 2攻击的防御方法。首先,我们通过对丽莎数据集上网络的第一层特征图进行频率分析来激发防御,这表明RP 2算法将高频噪声引入到输入图像中。为了去除高频噪声,我们在第一层之后引入了标准模糊核的深度卷积层。我们执行一个黑盒传输攻击,以表明低通过滤的特征图是更有益的比过滤输入。然后,我们提出了各种正则化方案,将这种低通滤波行为纳入网络的训练机制,并执行白盒攻击。我们得出结论,自适应攻击评估表明,攻击的成功率从90%下降到20%与总变差正则化,提出的防御之一。
Recently, the field of adversarial machine learning has been garnering attention by showing that state-of-the-art deep neural networks are vulnerable to adversarial examples, stemming from small perturbations being added to the input image. Adversarial examples are generated by a malicious adversary by obtaining access to the model parameters, such as gradient information, to alter the input or by attacking a substitute model and transferring those malicious examples over to attack the victim model. Specifically, one of these attack algorithms, Robust Physical Perturbations $(RP_{2})$, generates adversarial images of stop signs with black and white stickers to achieve high targeted misclassification rates against standard-architecture traffic sign classifiers. In this paper, we propose BlurNet, a defense against the RP2 attack. First, we motivate the defense with a frequency analysis of the first layer feature maps of the network on the LISA dataset, which shows that high frequency noise is introduced into the input image by the RP2 algorithm. To remove the high frequency noise, we introduce a depthwise convolution layer of standard blur kernels after the first layer. We perform a blackbox transfer attack to show that low-pass filtering the feature maps is more beneficial than filtering the input. We then present various regularization schemes to incorporate this low-pass filtering behavior into the training regime of the network and perform white-box attacks. We conclude with an adaptive attack evaluation to show that the success rate of the attack drops from 90% to 20% with total variation regularization, one of the proposed defenses.