Metric Learning for Adversarial Robustness

Metric Learning for Adversarial Robustness
复制标题

DOI:
--
复制
发表时间:
2019-09
影响因子:
6.8
通讯作者:
Chengzhi Mao;Ziyuan Zhong;Junfeng Yang;Carl Vondrick;Baishakhi Ray
Chengzhi Mao;Ziyuan Zhong;Junfeng Yang;Carl Vondrick;Baishakhi Ray
中科院分区:
医学1区
文献类型:
--
作者:
Chengzhi Mao;Ziyuan Zhong;Junfeng Yang;Carl Vondrick;Baishakhi Ray

文献摘要

被引文献

相似文献

众所周知,深层网络很容易受到对抗性攻击的影响。我们对最先进的 PGD 攻击方法下的深度表示进行了实证分析,发现该攻击导致内部表示向“错误”类别靠拢。受这一观察的启发,我们建议通过度量学习对受到攻击的表示空间进行正则化,以产生更强大的分类器。通过仔细采样度量学习的示例,我们学习到的表示不仅提高了鲁棒性,而且还可以检测到以前未见过的对抗性样本。定量实验表明,根据曲线下面积得分,与之前的工作相比,鲁棒性准确度提高了 4%,检测效率提高了 6%。我们的工作代码可在 https://github.com/columbia/Metric_Learning_Adversarial_Robustness 获取。
Deep networks are well-known to be fragile to adversarial attacks. We conduct an empirical analysis of deep representations under the state-of-the-art attack method called PGD, and find that the attack causes the internal representation to shift closer to the ``false'' class. Motivated by this observation, we propose to regularize the representation space under attack with metric learning to produce more robust classifiers. By carefully sampling examples for metric learning, our learned representation not only increases robustness, but also detects previously unseen adversarial samples. Quantitative experiments show improvement of robustness accuracy by up to 4% and detection efficiency by up to 6% according to Area Under Curve score over prior work. The code of our work is available at https://github.com/columbia/Metric_Learning_Adversarial_Robustness.