Tight Certification of Adversarially Trained Neural Networks via Nonconvex Low-Rank Semidefinite Relaxations

Tight Certification of Adversarially Trained Neural Networks via Nonconvex Low-Rank Semidefinite Relaxations
复制标题

DOI:
--
复制
发表时间:
2022-11
期刊:
--
影响因子:
--
通讯作者:
Hong-Ming Chiu;Richard Y. Zhang
Hong-Ming Chiu;Richard Y. Zhang
中科院分区:
其他
文献类型:
--
作者:
Hong-Ming Chiu;Richard Y. Zhang

文献摘要

相似文献

众所周知,对抗性训练可以产生高质量的神经网络模型,这些模型在经验上对对抗性扰动具有鲁棒性。然而,一旦对模型进行了对抗性训练,人们通常需要一个证明,证明该模型对所有未来的攻击都是真正健壮的。不幸的是,当面对对抗性训练的模型时,所有现有的方法都有很大的困难,无法获得足够强大的认证,从而在实践中发挥作用。线性规划(LP)技术尤其面临着一个“凸松弛障碍”,这使得它们无法进行高质量的认证,即使在使用混合整数线性规划(MILP)和分支定界(BnB)技术进行改进之后也是如此。本文提出了一种基于半定规划松弛的低秩约束的非凸证明技术。非凸松弛使得强认证可以与昂贵得多的SDP方法相媲美,而优化的变量却少得多,可以与弱得多的LP方法相媲美。尽管存在非凸性,但我们展示了如何使用现成的局部优化算法在多项式时间内实现和证明全局最优性。我们的实验发现,非凸松弛几乎完全消除了对对抗训练模型进行精确认证的差距。
Adversarial training is well-known to produce high-quality neural network models that are empirically robust against adversarial perturbations. Nevertheless, once a model has been adversarially trained, one often desires a certification that the model is truly robust against all future attacks. Unfortunately, when faced with adversarially trained models, all existing approaches have significant trouble making certifications that are strong enough to be practically useful. Linear programming (LP) techniques in particular face a"convex relaxation barrier"that prevent them from making high-quality certifications, even after refinement with mixed-integer linear programming (MILP) and branch-and-bound (BnB) techniques. In this paper, we propose a nonconvex certification technique, based on a low-rank restriction of a semidefinite programming (SDP) relaxation. The nonconvex relaxation makes strong certifications comparable to much more expensive SDP methods, while optimizing over dramatically fewer variables comparable to much weaker LP methods. Despite nonconvexity, we show how off-the-shelf local optimization algorithms can be used to achieve and to certify global optimality in polynomial time. Our experiments find that the nonconvex relaxation almost completely closes the gap towards exact certification of adversarially trained models.