Label-Noise Robust Domain Adaptation

Label-Noise Robust Domain Adaptation
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
Proceedings of machine learning research
影响因子:
--
通讯作者:
Xiyu Yu;Tongliang Liu;Mingming Gong;Kun Zhang;K. Batmanghelich;D. Tao
Xiyu Yu;Tongliang Liu;Mingming Gong;Kun Zhang;K. Batmanghelich;D. Tao
中科院分区:
其他
文献类型:
--
作者:
Xiyu Yu;Tongliang Liu;Mingming Gong;Kun Zhang;K. Batmanghelich;D. Tao

文献摘要

被引文献

相似文献

域适应的目的是在面临源(训练)和目标(测试)域之间的分布变化时纠正分类器。最先进的域适应方法利用深度网络来提取域不变表示。然而,现有的方法假设源域中的所有实例都被正确标记;而实际上,我们可能会获得带有噪声标签的源域,这并不奇怪。在本文中,我们首次全面研究了标签噪声如何在各种场景下对现有领域适应方法产生不利影响。此外,我们从理论上证明,存在一种可以从本质上减少域适应中噪声源标签的副作用的方法。具体来说,关注广义目标转移场景,其中标签分布 PY 和类条件分布 P X|Y 都可以改变,我们发现去噪条件不变分量(DCIC)框架可以证明确保(1)在给定源域中带有噪声标签的示例和目标域中未标记示例的情况下提取不变表示,以及(2)无偏差地估计目标域中的标签分布。合成数据和真实数据的实验结果验证了所提出方法的有效性。
Domain adaptation aims to correct the classifiers when faced with distribution shift between source (training) and target (test) domains. State-of-the-art domain adaptation methods make use of deep networks to extract domain-invariant representations. However, existing methods assume that all the instances in the source domain are correctly labeled; while in reality, it is unsurprising that we may obtain a source domain with noisy labels. In this paper, we are the first to comprehensively investigate how label noise could adversely affect existing domain adaptation methods in various scenarios. Further, we theoretically prove that there exists a method that can essentially reduce the side-effect of noisy source labels in domain adaptation. Specifically, focusing on the generalized target shift scenario, where both label distribution PY and the class-conditional distribution P X|Y can change, we discover that the denoising Conditional Invariant Component (DCIC) framework can provably ensures (1) extracting invariant representations given examples with noisy labels in the source domain and unlabeled examples in the target domain and (2) estimating the label distribution in the target domain with no bias. Experimental results on both synthetic and real-world data verify the effectiveness of the proposed method.