Implicit Bias of Gradient Descent based Adversarial Training on Separable Data

Implicit Bias of Gradient Descent based Adversarial Training on Separable Data
复制标题

DOI:
--
复制
发表时间:
2020-04
期刊:
--
影响因子:
--
通讯作者:
Yan Li;Ethan X. Fang;Huan Xu;T. Zhao
Yan Li;Ethan X. Fang;Huan Xu;T. Zhao
中科院分区:
其他
文献类型:
--
作者:
Yan Li;Ethan X. Fang;Huan Xu;T. Zhao

文献摘要

被引文献

相似文献

对抗训练是训练鲁棒神经网络的一种原则性方法。尽管在实践中取得了巨大的成功,但其理论特性仍然在很大程度上未被探索。在本文中,我们通过研究其计算特性,特别是其隐式偏差,为基于梯度下降的对抗训练提供了新的理论见解。我们以线性可分数据上的二进制分类任务为例,当参数沿沿着某些方向发散到无穷大时,损失渐近达到其下确界。具体来说,我们证明了对于任何固定的迭代T,当训练过程中的对抗扰动具有适当的有界L2范数时,通过基于梯度下降的对抗训练学习的分类器以O(1/\sqrt{T})$的速度收敛到最大L2范数边缘分类器,明显快于使用干净数据训练的速度O(1/\log T}$。此外,当训练期间的对抗性扰动具有有界Lq范数时,所得分类器在方向上收敛到最大混合范数边缘分类器,其具有鲁棒性的自然解释,作为在对数据的最坏情况有界Lq范数扰动下的最大L2范数边缘分类器。我们的研究结果为对抗训练提供了理论支持,它确实提高了对抗干扰的鲁棒性。
Adversarial training is a principled approach for training robust neural networks. Despite of tremendous successes in practice, its theoretical properties still remain largely unexplored. In this paper, we provide new theoretical insights of gradient descent based adversarial training by studying its computational properties, specifically on its implicit bias. We take the binary classification task on linearly separable data as an illustrative example, where the loss asymptotically attains its infimum as the parameter diverges to infinity along certain directions. Specifically, we show that for any fixed iteration $T$, when the adversarial perturbation during training has proper bounded L2 norm, the classifier learned by gradient descent based adversarial training converges in direction to the maximum L2 norm margin classifier at the rate of $O(1/\sqrt{T})$, significantly faster than the rate $O(1/\log T}$ of training with clean data. In addition, when the adversarial perturbation during training has bounded Lq norm, the resulting classifier converges in direction to a maximum mixed-norm margin classifier, which has a natural interpretation of robustness, as being the maximum L2 norm margin classifier under worst-case bounded Lq norm perturbation to the data. Our findings provide theoretical backups for adversarial training that it indeed promotes robustness against adversarial perturbation.