Improved techniques for deterministic l2 robustness

Improved techniques for deterministic l2 robustness
复制标题

DOI:
10.48550/arxiv.2211.08453
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Sahil Singla;S. Feizi
Sahil Singla;S. Feizi
中科院分区:
其他
文献类型:
--
作者:
Sahil Singla;S. Feizi

文献摘要

相似文献

在 $l_{2}$ 范数下训练具有严格 1-Lipschitz 约束的卷积神经网络 (CNN) 对于对抗鲁棒性、可解释梯度和稳定训练非常有用。 1-Lipschitz CNN 的设计通常是强制每一层都具有正交雅可比矩阵(对于所有输入),以防止梯度在反向传播过程中消失。然而,它们的性能通常明显落后于强制实施 Lipschitz 约束的启发式方法,其中生成的 CNN 不是\textit{可证明} 1-Lipschitz。在这项工作中,我们通过引入(a)一种通过用 1-隐藏层 MLP 替换最后一个线性层来验证 1-Lipschitz CNN 鲁棒性的过程来缩小这一差距,该过程显着提高了标准精度和可证明鲁棒精度的性能,(b)一种显着减少斜正交卷积(SOC)层每轮训练时间的方法(对于更深的网络减少了 30%以上)和(c)一类使用数学属性的池化层输入到流形的 $l_{2}$ 距离是 1-Lipschitz。使用这些方法,我们显着提高了 CIFAR-10 上最先进的标准和可证明的鲁棒精度(增益 +1.79\% 和 +3.82\%),以及 CIFAR-100 上的类似精度(+3.78\% 和 +4.75\%)。代码可在 \url{https://github.com/singlasahil14/improved_l2_robustness} 获取。
Training convolutional neural networks (CNNs) with a strict 1-Lipschitz constraint under the $l_{2}$ norm is useful for adversarial robustness, interpretable gradients and stable training. 1-Lipschitz CNNs are usually designed by enforcing each layer to have an orthogonal Jacobian matrix (for all inputs) to prevent the gradients from vanishing during backpropagation. However, their performance often significantly lags behind that of heuristic methods to enforce Lipschitz constraints where the resulting CNN is not \textit{provably} 1-Lipschitz. In this work, we reduce this gap by introducing (a) a procedure to certify robustness of 1-Lipschitz CNNs by replacing the last linear layer with a 1-hidden layer MLP that significantly improves their performance for both standard and provably robust accuracy, (b) a method to significantly reduce the training time per epoch for Skew Orthogonal Convolution (SOC) layers (>30\% reduction for deeper networks) and (c) a class of pooling layers using the mathematical property that the $l_{2}$ distance of an input to a manifold is 1-Lipschitz. Using these methods, we significantly advance the state-of-the-art for standard and provable robust accuracies on CIFAR-10 (gains of +1.79\% and +3.82\%) and similarly on CIFAR-100 (+3.78\% and +4.75\%) across all networks. Code is available at \url{https://github.com/singlasahil14/improved_l2_robustness}.