Hitting Depth : Investigating Robustness to Adversarial Examples in Deep Convolutional Neural Networks

Hitting Depth : Investigating Robustness to Adversarial Examples in Deep Convolutional Neural Networks
复制标题

命中深度:研究深度卷积神经网络中对抗性示例的鲁棒性

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
C. Billovits
C. Billovits
中科院分区:
--
文献类型:
--
作者:
C. Billovits

文献摘要

被引文献

相似文献

包括卷积神经网络(CNN)在内的机器学习模型很容易受到敌意输入图像的影响,这些图像被干扰,故意愚弄模型到一个错误的、高置信度的预测中,而不是视觉上可感知的变化。以前的工作表明,高维线性导致了这些对抗性空间。我们首先使用VGGNet验证关于基于梯度和基于模式的对抗性例子的泛化的假设。我们展示了一个过程,用于可视化和识别对抗性图像和其常规对应图像之间的激活变化。最后,我们在一个新的贝叶斯框架中利用这两种方法的信息来增强对敌意例子的L2稳健性。利用该框架,我们成功地提高了对抗性实例的预测精度。
Machine learning models, including Convolutional Neural Networks (CNN) are susceptible to adversarial examples input images that have been perturbed to deliberately fool a model into an incorrect, high-confidence prediction without a visually perceptible change. Previous work shows that high-dimensional linearities cause these adversarial pockets of space. We first validate assumptions about the generalization of gradient-based and pattern-based adversarial examples using VGGNet. We show a process for visualizing and identifying changes in activations between adversarial images and their regular counterparts. Finally, we leverage information from these two approaches in a novel Bayesian framework to increase l2 robustness to adversarial examples. Using this framework, we successfully improve the prediction accuracy on adversarial examples.