Robustness May Be at Odds with Accuracy

Robustness May Be at Odds with Accuracy
复制标题

DOI:
--
复制
发表时间:
2018-05
期刊:
arXiv: Machine Learning
影响因子:
--
通讯作者:
Dimitris Tsipras;Shibani Santurkar;Logan Engstrom;Alexander Turner;A. Madry
Dimitris Tsipras;Shibani Santurkar;Logan Engstrom;Alexander Turner;A. Madry
中科院分区:
其他
文献类型:
--
作者:
Dimitris Tsipras;Shibani Santurkar;Logan Engstrom;Alexander Turner;A. Madry

文献摘要

被引文献

相似文献

我们表明,对抗鲁棒性的目标与标准概括的目标之间可能存在固有的张力。具体而言,训练健壮的模型不仅可能更耗资资源,而且还会导致标准准确性的降低。我们证明,模型的标准准确性与对抗性扰动的鲁棒性之间的这种权衡证明存在于相当简单和自然的环境中。这些发现也证实了在更复杂的环境中经验观察到的类似现象。此外,我们认为这种现象是强大的分类器与标准分类器学习根本不同的特征表示形式的结果。尤其是这些差异似乎会带来意外的好处:强大模型所学的表示倾向于与显着的数据特征和人类看法更好地保持一致。
We show that there may exist an inherent tension between the goal of adversarial robustness and that of standard generalization. Specifically, training robust models may not only be more resource-consuming, but also lead to a reduction of standard accuracy. We demonstrate that this trade-off between the standard accuracy of a model and its robustness to adversarial perturbations provably exists in a fairly simple and natural setting. These findings also corroborate a similar phenomenon observed empirically in more complex settings. Further, we argue that this phenomenon is a consequence of robust classifiers learning fundamentally different feature representations than standard classifiers. These differences, in particular, seem to result in unexpected benefits: the representations learned by robust models tend to align better with salient data characteristics and human perception.