Indirect Invisible Poisoning Attacks on Domain Adaptation

Indirect Invisible Poisoning Attacks on Domain Adaptation
复制标题

DOI:
10.1145/3447548.3467214
复制
发表时间:
2021-08
期刊:
Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Jun Wu;Jingrui He
Jun Wu;Jingrui He
中科院分区:
其他
文献类型:
--
作者:
Jun Wu;Jingrui He

文献摘要

被引文献

相似文献

无监督域适应已成功应用于多个具有重大影响的应用中,因为当源域和目标域相关时,它提高了学习算法的泛化性能。然而,域适应模型的对抗脆弱性在很大程度上被忽视了。大多数现有的无监督域适应算法可能很容易被对手愚弄,当从被恶意操纵的源域转移知识时,会导致在目标域上的预测性能下降。为了证明现有域适应技术的对抗脆弱性,在本文中,我们提出了一个通用的数据投毒攻击框架,名为I2Attack用于域适应,它具有以下特性:(1) 难以察觉:所有被投毒的输入看起来都很自然;(2) 间接对抗:只有源样本被恶意操纵;(3) 算法上不可见:源分类错误以及源域和目标域之间的边缘域差异都不会增加。具体而言,它旨在通过最大化源域和目标域在输入特征空间和类别标签空间上的有标签域差异来降低目标域上的整体预测性能。在这个框架内,提出了一系列实用的投毒攻击来愚弄与不同差异度量相关的现有域适应算法。在各种域适应基准上的大量实验证实了我们提出的I2Attack框架的有效性和计算效率。
Unsupervised domain adaptation has been successfully applied across multiple high-impact applications, since it improves the generalization performance of a learning algorithm when the source and target domains are related. However, the adversarial vulnerability of domain adaptation models has largely been neglected. Most existing unsupervised domain adaptation algorithms might be easily fooled by an adversary, resulting in deteriorated prediction performance on the target domain, when transferring the knowledge from a maliciously manipulated source domain. To demonstrate the adversarial vulnerability of existing domain adaptation techniques, in this paper, we propose a generic data poisoning attack framework named I2Attack for domain adaptation with the following properties: (1) perceptibly unnoticeable: all the poisoned inputs are natural-looking; (2)adversarially indirect: only source examples are maliciously manipulated; (3) algorithmically invisible: both source classification error and marginal domain discrepancy between source and target domains will not increase. Specifically, it aims to degrade the overall prediction performance on the target domain by maximizing the label-informed domain discrepancy over both input feature space and class-label space be-tween source and target domains. Within this framework, a family of practical poisoning attacks are presented to fool the existing domain adaptation algorithms associated with different discrepancy measures. Extensive experiments on various domain adaptation benchmarks confirm the effectiveness and computational efficiency of our proposed I2Attack framework.