Breaking certified defenses: Semantic adversarial examples with spoofed robustness certificates

Breaking certified defenses: Semantic adversarial examples with spoofed robustness certificates
复制标题

DOI:
--
复制
发表时间:
2020-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Amin Ghiasi;Ali Shafahi;T. Goldstein
Amin Ghiasi;Ali Shafahi;T. Goldstein
中科院分区:
其他
文献类型:
--
作者:
Amin Ghiasi;Ali Shafahi;T. Goldstein

文献摘要

相似文献

对抗性攻击的防御可以分为认证和非认证两种。可认证防御使网络在一定的$\ell_p$有界半径内具有鲁棒性,因此攻击者不可能在证书范围内制作对抗性示例。我们提出了一种在认证半径之外保持对抗性样本的不可感知性的攻击方法。此外,提出的“影子攻击”可以通过产生一个难以察觉的对抗示例来欺骗可认证的健壮网络,该示例被错误分类并产生一个强大的“欺骗”证书。
Defenses against adversarial attacks can be classified into certified and non-certified. Certifiable defenses make networks robust within a certain $\ell_p$-bounded radius, so that it is impossible for the adversary to make adversarial examples in the certificate bound. We present an attack that maintains the imperceptibility property of adversarial examples while being outside of the certified radius. Furthermore, the proposed "Shadow Attack" can fool certifiably robust networks by producing an imperceptible adversarial example that gets misclassified and produces a strong ``spoofed'' certificate.