Distributed Adversarial Training to Robustify Deep Neural Networks at Scale

Distributed Adversarial Training to Robustify Deep Neural Networks at Scale
复制标题

DOI:
10.48550/arxiv.2206.06257
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Gaoyuan Zhang;Songtao Lu;Yihua Zhang;Xiangyi Chen;Pin-Yu Chen;Quanfu Fan;Lee Martie;L. Horesh
Gaoyuan Zhang;Songtao Lu;Yihua Zhang;Xiangyi Chen;Pin-Yu Chen;Quanfu Fan;Lee Martie;L. Horesh
中科院分区:
其他
文献类型:
--
作者:
Gaoyuan Zhang;Songtao Lu;Yihua Zhang;Xiangyi Chen;Pin-Yu Chen;Quanfu Fan;Lee Martie;L. Horesh

文献摘要

相似文献

当前的深度神经网络(DNN)很容易受到敌意攻击,对输入的敌意扰动可以改变或操纵分类。为了防御这种攻击,一种被称为对抗性训练(AT)的有效和流行的方法已经被证明通过最小-最大稳健训练方法来减轻对抗性攻击的负面影响。虽然有效,但它是否能成功地适应分布式学习环境仍不清楚。在多台机器上进行分布式优化的能力使我们能够在大型模型和数据集上扩大健壮的训练。受此启发,我们提出了分布式对抗训练(DAT),这是一种在多台机器上实现的大批量对抗训练框架。我们证明DAT是通用的,它支持对有标签和无标签数据的训练,支持多种类型的攻击生成方法,以及有利于分布式优化的梯度压缩操作。理论上,在最优化理论的标准条件下,我们给出了一般非凸集上DAT收敛到一阶驻点的收敛速度。在实验上,我们证明了DAT匹配或超过了最先进的稳健精度,并实现了优雅的训练加速比(例如,在ImageNet下的ResNet-50上)。有关代码,请访问https://github.com/dat-2022/dat.
Current deep neural networks (DNNs) are vulnerable to adversarial attacks, where adversarial perturbations to the inputs can change or manipulate classification. To defend against such attacks, an effective and popular approach, known as adversarial training (AT), has been shown to mitigate the negative impact of adversarial attacks by virtue of a min-max robust training method. While effective, it remains unclear whether it can successfully be adapted to the distributed learning context. The power of distributed optimization over multiple machines enables us to scale up robust training over large models and datasets. Spurred by that, we propose distributed adversarial training (DAT), a large-batch adversarial training framework implemented over multiple machines. We show that DAT is general, which supports training over labeled and unlabeled data, multiple types of attack generation methods, and gradient compression operations favored for distributed optimization. Theoretically, we provide, under standard conditions in the optimization theory, the convergence rate of DAT to the first-order stationary points in general non-convex settings. Empirically, we demonstrate that DAT either matches or outperforms state-of-the-art robust accuracies and achieves a graceful training speedup (e.g., on ResNet-50 under ImageNet). Codes are available at https://github.com/dat-2022/dat.