Towards Adversarial Attack Resistant Deep Neural Networks

Towards Adversarial Attack Resistant Deep Neural Networks
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Tiago A. O. Alves;S. Kundu
Tiago A. O. Alves;S. Kundu
中科院分区:
其他
文献类型:
--
作者:
Tiago A. O. Alves;S. Kundu

文献摘要

相似文献

。最近的出版物表明,基于神经网络的分类器很容易受到对抗性输入的影响,这些输入实际上与正常数据没有区别,而这些数据是为了强制错误分类而明确构建的。在本文中,我们提出了几种应对这些威胁的防御措施。首先,我们观察到大多数对抗性攻击通过在模型返回的置信度上安装梯度上升来成功,这使得对手能够了解分类边界。我们的防御基于拒绝访问精确的分类边界。我们的第一道防御在输出置信水平中添加了受控随机噪声,这可以防止对手收敛其数值近似攻击。我们的下一个防御是基于这样的观察:通过改变训练顺序,我们通常会得到提供相同分类精度的模型,但它们在数值上不同。这些模型的集合允许我们在查询期间在这些等效模型之间随机切换,这进一步模糊了分类边界。我们通过对抗性输入生成器展示了我们的防御,该生成器击败了先前发布的防御,但不能突破所提出的防御对其非静态性质的影响。
. Recent publications have shown that neural network based classifiers are vulnerable to adversarial inputs that are virtually indistinguishable from normal data, constructed explicitly for the purpose of forcing misclassification. In this paper, we present several defenses to counter these threats. First, we observe that most adversarial attacks succeed by mounting gradient ascent on the confidence returned by the model, which allows adversary to gain understanding of the classification boundary. Our defenses are based on denying access to the precise classification boundary. Our first defense adds a controlled random noise to the output confidence levels, which prevents an adversary from converging in their numerical approximation attack. Our next defense is based on the observation that by varying the order of the training, often we arrive at models which offer the same classification accuracy, yet they are different numerically. An ensemble of such models allows us to randomly switch between these equivalent models during query which further blurs the classification boundary. We demonstrate our defense via an adversarial input generator which defeats previously published defenses but cannot breach the proposed defenses do to their non-static nature.