SAT: Improving Adversarial Training via Curriculum-Based Loss Smoothing

SAT: Improving Adversarial Training via Curriculum-Based Loss Smoothing
复制标题

DOI:
10.1145/3474369.3486878
复制
发表时间:
2020-03
期刊:
Proceedings of the 14th ACM Workshop on Artificial Intelligence and Security
影响因子:
--
通讯作者:
Chawin Sitawarin;S. Chakraborty;David A. Wagner
Chawin Sitawarin;S. Chakraborty;David A. Wagner
中科院分区:
其他
文献类型:
--
作者:
Chawin Sitawarin;S. Chakraborty;David A. Wagner

文献摘要

相似文献

对抗训练(AT)已成为训练健壮网络的一种流行选择。然而,它往往牺牲干净的精度以牺牲稳健性,并遭受较大的泛化误差。为了解决这些问题,我们提出了平滑对抗性训练(SAT),以我们对损失Hessian的特征谱的分析为指导。我们发现,课程学习,一种强调开始“容易”并逐渐增加训练的“难度”的计划,为适当选择的难度度量平滑了对抗性失败的图景。我们给出了对抗性环境下课程学习的一般公式,并提出了两种基于最大Hessian特征值(H-SAT)和Softmax概率(P-SA)的难度度量。我们证明,与AT相比,SAT即使在很大的扰动范数下也能稳定网络训练,并允许网络以更干净的精度与稳健性权衡曲线运行。这导致了与AT、TRADS和其他基准相比,在干净的准确性和健壮性方面都有了显著的改进。为了突出一些结果,我们最好的模型在CIFAR-100上比AT分别提高了6%和1%的正常和稳健精度。在ImageNet的十类子集Imagenette上,我们的模型在正常精度和稳健精度上分别比AT高23%和3%。
Adversarial training (AT) has become a popular choice for training robust networks. However, it tends to sacrifice clean accuracy heavily in favor of robustness and suffers from a large generalization error. To address these concerns, we propose Smooth Adversarial Training (SAT), guided by our analysis on the eigenspectrum of the loss Hessian. We find that curriculum learning, a scheme that emphasizes on starting "easy'' and gradually ramping up on the "difficulty'' of training, smooths the adversarial loss landscape for a suitably chosen difficulty metric. We present a general formulation for curriculum learning in the adversarial setting and propose two difficulty metrics based on the maximal Hessian eigenvalue (H-SAT) and the softmax probability (P-SA). We demonstrate that SAT stabilizes network training even for a large perturbation norm and allows the network to operate at a better clean accuracy versus robustness trade-off curve compared to AT. This leads to a significant improvement in both clean accuracy and robustness compared to AT, TRADES, and other baselines. To highlight a few results, our best model improves normal and robust accuracy by 6% and 1% on CIFAR-100 compared to AT, respectively. On Imagenette, a ten-class subset of ImageNet, our model outperforms AT by 23% and 3% on normal and robust accuracy respectively.