Learning to Defend by Learning to Attack

Learning to Defend by Learning to Attack
复制标题

DOI:
--
复制
发表时间:
2018-11
期刊:
--
影响因子:
--
通讯作者:
Haoming Jiang;Zhehui Chen;Yuyang Shi;Bo Dai;T. Zhao
Haoming Jiang;Zhehui Chen;Yuyang Shi;Bo Dai;T. Zhao
中科院分区:
其他
文献类型:
--
作者:
Haoming Jiang;Zhehui Chen;Yuyang Shi;Bo Dai;T. Zhao

文献摘要

相似文献

对抗训练为训练鲁棒神经网络提供了一种原则性方法。从优化的角度来看,对抗训练本质上是解决一个双层优化问题。领导者问题试图学习一个鲁棒的分类器,而追随者问题试图生成对抗样本。不幸的是,这样的双层问题是很难解决,由于其高度复杂的结构。这项工作提出了一种新的基于通用学习学习(L2 L)框架的对抗训练方法。具体来说,我们不是将现有的手工设计的算法应用于内部问题,而是学习一个优化器,它被参数化为卷积神经网络。同时,学习一个鲁棒的分类器来防御由学习的优化器产生的对抗性攻击。在CIFAR-10和CIFAR-100数据集上的实验表明,L2 L在分类精度和计算效率方面都优于现有的对抗训练方法。此外,我们的L2 L框架可以扩展到生成对抗模仿学习和稳定的训练。
Adversarial training provides a principled approach for training robust neural networks. From an optimization perspective, adversarial training is essentially solving a bilevel optimization problem. The leader problem is trying to learn a robust classifier, while the follower problem is trying to generate adversarial samples. Unfortunately, such a bilevel problem is difficult to solve due to its highly complicated structure. This work proposes a new adversarial training method based on a generic learning-to-learn (L2L) framework. Specifically, instead of applying existing hand-designed algorithms for the inner problem, we learn an optimizer, which is parametrized as a convolutional neural network. At the same time, a robust classifier is learned to defense the adversarial attack generated by the learned optimizer. Experiments over CIFAR-10 and CIFAR-100 datasets demonstrate that L2L outperforms existing adversarial training methods in both classification accuracy and computational efficiency. Moreover, our L2L framework can be extended to generative adversarial imitation learning and stabilize the training.