Toward Adversarial Robustness in Unlabeled Target Domains

Toward Adversarial Robustness in Unlabeled Target Domains
复制标题

DOI:
10.1109/tip.2023.3242141
复制
发表时间:
2023-02
影响因子:
10.6
通讯作者:
Jiajin Zhang;Hanqing Chao;Pingkun Yan
Jiajin Zhang;Hanqing Chao;Pingkun Yan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jiajin Zhang;Hanqing Chao;Pingkun Yan

文献摘要

相似文献

在过去的几年中,人们发明了各种对抗训练(AT)方法来增强深度学习模型对抗对抗攻击的能力。然而,主流AT方法假设训练和测试数据来自相同的分布,并且训练数据被注释。当这两个假设被违反时,现有的 AT 方法就会失败,因为它们要么无法将从源域学到的知识传递到未标记的目标域,要么被该未标记空间中的对抗样本混淆。在本文中,我们首先指出这个新的且具有挑战性的问题——未标记目标域中的对抗性训练。然后,我们提出了一种名为无监督跨域对抗训练(UCAT)的新颖框架来解决这个问题。 UCAT 在自动选择的未注释目标域数据的高质量伪标签以及源域数据的判别性和鲁棒性锚表示的指导下,有效地利用标记源域的知识来防止对抗性样本误导训练过程。在四个公共基准上的实验表明,使用UCAT训练的模型可以实现高精度和强鲁棒性。所提出的组件的有效性通过大量的消融研究得到了证明。源代码可在 https://github.com/DIAL-RPI/UCAT 上公开获取。
In the past several years, various adversarial training (AT) approaches have been invented to robustify deep learning model against adversarial attacks. However, mainstream AT methods assume the training and testing data are drawn from the same distribution and the training data are annotated. When the two assumptions are violated, existing AT methods fail because either they cannot pass knowledge learnt from a source domain to an unlabeled target domain or they are confused by the adversarial samples in that unlabeled space. In this paper, we first point out this new and challenging problem— adversarial training in unlabeled target domain. We then propose a novel framework named Unsupervised Cross-domain Adversarial Training (UCAT) to address this problem. UCAT effectively leverages the knowledge of the labeled source domain to prevent the adversarial samples from misleading the training process, under the guidance of automatically selected high quality pseudo labels of the unannotated target domain data together with the discriminative and robust anchor representations of the source domain data. The experiments on four public benchmarks show that models trained with UCAT can achieve both high accuracy and strong robustness. The effectiveness of the proposed components is demonstrated through a large set of ablation studies. The source code is publicly available at https://github.com/DIAL-RPI/UCAT.