Exploiting the Inherent Limitation of L0 Adversarial Examples

Exploiting the Inherent Limitation of L0 Adversarial Examples
复制标题

DOI:
--
复制
发表时间:
2018-12
期刊:
--
影响因子:
--
通讯作者:
F. Zuo;Bokai Yang;Xiaopeng Li;Qiang Zeng
F. Zuo;Bokai Yang;Xiaopeng Li;Qiang Zeng
中科院分区:
其他
文献类型:
--
作者:
F. Zuo;Bokai Yang;Xiaopeng Li;Qiang Zeng

文献摘要

被引文献

相似文献

尽管神经网络在图像分类等任务上取得了巨大成就,但它们很脆弱,容易受到对抗性示例(AE)攻击的攻击,这种攻击是通过在输入中添加人类难以察觉的扰动来制作的,从而使基于神经网络的分类器错误地标记它们。特别是,L0 ae是一类被广泛讨论的威胁,其中攻击者可以破坏的像素数量受到限制。然而,我们的观察是,虽然L0攻击修改尽可能少的像素,但它们往往会对修改的像素造成大幅度的扰动。我们认为这是L0 AEs的固有限制,并通过检测和纠正它们来阻止此类攻击。该检测器的主要新颖之处在于,我们利用L0攻击的固有局限性,将声发射检测问题转化为比较问题。更具体地说,给定图像I,对其进行预处理以获得另一个图像I'。已知比较有效的Siamese网络将I和I’作为输入对来确定I是否为AE。经过训练的暹罗网络会自动准确地捕捉I和I'之间的差异,以检测L0扰动。此外,我们表明,用于检测的预处理技术inpainting也可以作为一种有效的防御,它有很高的概率消除L0扰动的不利影响。因此,我们的系统AEPECKER不仅具有较高的声发射检测精度,而且具有显著的校正分类结果的能力。
Despite the great achievements made by neural networks on tasks such as image classification, they are brittle and vulnerable to adversarial example (AE) attacks, which are crafted by adding human-imperceptible perturbations to inputs in order that a neural-network-based classifier incorrectly labels them. In particular, L0 AEs are a category of widely discussed threats where adversaries are restricted in the number of pixels that they can corrupt. However, our observation is that, while L0 attacks modify as few pixels as possible, they tend to cause large-amplitude perturbations to the modified pixels. We consider this as an inherent limitation of L0 AEs, and thwart such attacks by both detecting and rectifying them. The main novelty of the proposed detector is that we convert the AE detection problem into a comparison problem by exploiting the inherent limitation of L0 attacks. More concretely, given an image I, it is pre-processed to obtain another image I' . A Siamese network, which is known to be effective in comparison, takes I and I' as the input pair to determine whether I is an AE. A trained Siamese network automatically and precisely captures the discrepancies between I and I' to detect L0 perturbations. In addition, we show that the pre-processing technique, inpainting, used for detection can also work as an effective defense, which has a high probability of removing the adversarial influence of L0 perturbations. Thus, our system, called AEPECKER, demonstrates not only high AE detection accuracies, but also a notable capability to correct the classification results.