Unsupervised domain adaptation based on cluster matching and Fisher criterion for image classification

Unsupervised domain adaptation based on cluster matching and Fisher criterion for image classification
复制标题

基于聚类匹配和 Fisher 准则的图像分类无监督域自适应

DOI:
10.1016/j.compeleceng.2021.107041
复制
发表时间:
2021
影响因子:
4.3
通讯作者:
Chen Yang
Chen Yang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chang Heyou;Zhang Fanlong;Ma Shuai;Gao Guangwei;Zheng Hao;Chen Yang

文献摘要

被引文献

相似文献

当两个域具有不同的分布时,将从标记域(源域)学习到的知识转移到未标记域(目标域)是具有挑战性的。解决这一问题的关键是减小两个域之间的分布漂移。为了对齐分布,大多数现有的工作首先学习源域上的分类器以获得目标样本的伪标签,然后基于伪标签计算目标域分布。然而,分类器可能不满足目标域,因为它在学习过程中忽略了目标分布。错误标记的样本将在目标域分布的计算中引起较大的误差。为了解决这个问题,我们提出了一种新的方法,命名为聚类匹配和Fisher准则(CMFC),生成一个准确的伪标签,为每个目标样本在一个潜在的歧视性子空间,考虑两个域的分布。具体地说,我们首先在潜在子空间中分别对两个域中的样本进行聚类,然后将目标域中的聚类质心与源域中的类质心进行匹配。通过集群匹配考虑这两个域分布,以分配更准确的伪标签。此外,我们利用Fisher准则来最小化类内方差,同时最大化类间方差,这有利于进一步减少分布偏移。我们将聚类匹配和Fisher准则结合到一个统一的模型中,并设计了一个ADMM算法来有效地解决所提出的方法。在五个数据集上的分类实验表明了CMFC的优越性。
Transferring knowledge learned from a labeled domain (source domain) to an unlabeled domain (target domain) is challenging when the two domains have different distributions. The key to the problem is to reduce the distribution shift between the two domains. To align the distributions, most existing works first learn a classifier on the source domain to obtain pseud-labels for target samples, then calculate the target domain distribution based on the pseud-labels. However, the classifier may not meet the target domain because it loses sight of the target distribution during the learning procedure. The mislabeled samples will cause large errors in the calculation of the target domain distribution. To address this issue, we propose a novel method, named cluster matching and Fisher criterion (CMFC), to generate an accurate pseudo-label for each target sample in a latent discriminative subspace by considering both domain distributions. Specifically, we first cluster the samples in both domains, respectively, in the latent subspace and then match the cluster centroid in the target domain with the class centroid in the source domain. Both domain distributions are taken into consideration via cluster matching to assign more accurate pseud-labels. Moreover, we leverage the Fisher criterion to minimize intra-class variances while maximizing inter-class variances, which is conducive to further reducing the distribution shift. We incorporate cluster matching and the Fisher criterion into a united model and design an ADMM algorithm to effectively solve the proposed method. Extensive experiments on five datasets for classification tasks demonstrate the superiority of CMFC.