On the Algorithmic Stability of Adversarial Training

On the Algorithmic Stability of Adversarial Training
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Yue Xing;Qifan Song;Guang Cheng
Yue Xing;Qifan Song;Guang Cheng
中科院分区:
其他
文献类型:
--
作者:
Yue Xing;Qifan Song;Guang Cheng

文献摘要

被引文献

相似文献

对抗性训练是弥补深度学习模型对抗对抗性攻击脆弱性的流行工具,关于对抗性训练算法的训练损失有丰富的理论文献。相比之下,本文研究了通用对抗训练算法的算法稳定性,这可以进一步帮助建立泛化误差的上限。通过计算出稳定性上限和下限,我们认为对抗性训练的不可微性问题会导致比自然算法更差的算法稳定性。为了解决这个问题,我们考虑噪声注入方法。虽然不可微性问题严重影响对抗训练的稳定性,但注入噪声可以使训练轨迹避免出现占主导地位的不可微性问题,从而增强对抗训练的稳定性能。我们的分析还研究了算法稳定性与对抗性攻击的数值逼近误差之间的关系。
The adversarial training is a popular tool to remedy the vulnerability of deep learning models against adversarial attacks, and there is rich theoretical literature on the training loss of adversarial training algorithms. In contrast, this paper studies the algorithmic stability of a generic adversarial training algorithm, which can further help to establish an upper bound for generalization error. By figuring out the stability upper bound and lower bound, we argue that the non-differentiability issue of adversarial training causes worse algorithmic stability than their natu-ral counterparts. To tackle this problem, we consider a noise injection method. While the non-differentiability problem seriously affects the stability of adversarial training, injecting noise enables the training trajectory to avoid the occurrence of non-differentiability with dominating probability, hence enhancing the stability performance of adversarial training. Our analysis also studies the relation between the algorithm stability and numerical approximation error of adversarial attacks.