Constrained Gradient Descent: A Powerful and Principled Evasion Attack Against Neural Networks

Constrained Gradient Descent: A Powerful and Principled Evasion Attack Against Neural Networks
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
--
影响因子:
--
通讯作者:
Weiran Lin;Keane Lucas;Lujo Bauer;M. Reiter;Mahmood Sharif
Weiran Lin;Keane Lucas;Lujo Bauer;M. Reiter;Mahmood Sharif
中科院分区:
其他
文献类型:
--
作者:
Weiran Lin;Keane Lucas;Lujo Bauer;M. Reiter;Mahmood Sharif

文献摘要

被引文献

相似文献

我们提出了针对深层神经网络的新的,更有效的目标白盒攻击。我们的攻击可以更好地与攻击者的目标保持一致:(1)欺骗模型将比其他类别分配给目标类别的概率更高,而(2)保持攻击输入的$ \ epsilon $ distance。首先,我们演示了一个明确编码(1)的损耗函数,并证明自动PGD可以通过它找到更多攻击。其次,我们使用捕获(1)和(2)的损耗函数的改进,提出了一种新的攻击方法,即约束梯度下降(CGD)。 CGD试图满足这两个攻击者目标 - 错误分类和有限的$ \ ell_ {p} $ - 标准 - 作为优化的一部分,而不是通过临时的后处理技术(例如,投影或剪辑) 。我们表明,在CIFAR10(0.9--4.2%)和Imagenet(8.6--13.6%)上,CGD比最先进的攻击更为成功,同时消耗的时间更少(11.4---18.8%)。统计测试证实,我们的攻击优于其他人在不同数据集和$ \ epsilon $的值上的领先辩护。
We propose new, more efficient targeted white-box attacks against deep neural networks. Our attacks better align with the attacker's goal: (1) tricking a model to assign higher probability to the target class than to any other class, while (2) staying within an $\epsilon$-distance of the attacked input. First, we demonstrate a loss function that explicitly encodes (1) and show that Auto-PGD finds more attacks with it. Second, we propose a new attack method, Constrained Gradient Descent (CGD), using a refinement of our loss function that captures both (1) and (2). CGD seeks to satisfy both attacker objectives -- misclassification and bounded $\ell_{p}$-norm -- in a principled manner, as part of the optimization, instead of via ad hoc post-processing techniques (e.g., projection or clipping). We show that CGD is more successful on CIFAR10 (0.9--4.2%) and ImageNet (8.6--13.6%) than state-of-the-art attacks while consuming less time (11.4--18.8%). Statistical tests confirm that our attack outperforms others against leading defenses on different datasets and values of $\epsilon$.