Adversarial Training Can Hurt Generalization

Adversarial Training Can Hurt Generalization
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Aditi Raghunathan;Sang Michael Xie;Fanny Yang;John C. Duchi;Percy Liang
Aditi Raghunathan;Sang Michael Xie;Fanny Yang;John C. Duchi;Percy Liang
中科院分区:
其他
文献类型:
--
作者:
Aditi Raghunathan;Sang Michael Xie;Fanny Yang;John C. Duchi;Percy Liang

文献摘要

相似文献

尽管对抗训练可以提高稳健的精度(针对对手),但有时会损害标准精度(当没有对手时)。先前的工作已经研究了标准精度和稳健精度之间的这种权衡,但仅在无预测器在无限数据限制中的两个目标上都表现良好的情况下。在本文中,我们表明,即使具有无限数据的最佳预测变量在两个目标上都效果很好,但折衷仍然可以用有限的数据表现出来。此外,由于我们的构建是基于凸学习问题的,因此我们排除了优化问题,因此在稳健性和概括之间产生了根本的张力。最后,我们表明,强大的自我训练主要通过利用未标记的数据来消除这种权衡。
While adversarial training can improve robust accuracy (against an adversary), it sometimes hurts standard accuracy (when there is no adversary). Previous work has studied this tradeoff between standard and robust accuracy, but only in the setting where no predictor performs well on both objectives in the infinite data limit. In this paper, we show that even when the optimal predictor with infinite data performs well on both objectives, a tradeoff can still manifest itself with finite data. Furthermore, since our construction is based on a convex learning problem, we rule out optimization concerns, thus laying bare a fundamental tension between robustness and generalization. Finally, we show that robust self-training mostly eliminates this tradeoff by leveraging unlabeled data.