Distributional Shift Adaptation using Domain-Specific Features

Distributional Shift Adaptation using Domain-Specific Features
复制标题

DOI:
10.1109/bigdata55660.2022.10020444
复制
发表时间:
2022-11
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Anique Tahir;Lu Cheng;Ruocheng Guo;Huan Liu
Anique Tahir;Lu Cheng;Ruocheng Guo;Huan Liu
中科院分区:
其他
文献类型:
--
作者:
Anique Tahir;Lu Cheng;Ruocheng Guo;Huan Liu

文献摘要

相似文献

机器学习算法通常假设训练样本和测试样本来自相同的分布,即内分布。然而,在开放世界的场景中,大数据流传输(OOD)可能会导致这些算法效率低下。OOD挑战的现有解决方案试图识别不同训练域中的不变特征。基本的假设是,这些不变特征也应该在未标记的目标域中工作得相当好。相比之下,这项工作对特定于领域的特征感兴趣,这些特征既包括不变特征,也包括目标领域特有的特征。我们提出了一种简单而有效的方法,该方法一般依赖于相关性,而不考虑特征是否不变。我们的方法使用OOD基本模型(教师模型)识别的最有把握的预测样本来训练一个有效适应目标领域的新模型(学生模型)。在基准数据集上的经验评估表明,该算法的性能比SOTA提高了10%-20%。
Machine learning algorithms typically assume that the training and test samples come from the same distributions, i.e., in-distribution. However, in open-world scenarios, streaming big data can be Out-Of-Distribution (OOD), rendering these algorithms ineffective. Prior solutions to the OOD challenge seek to identify invariant features across different training domains. The underlying assumption is that these invariant features should also work reasonably well in the unlabeled target domain. By contrast, this work is interested in the domain-specific features that include both invariant features and features unique to the target domain. We propose a simple yet effective approach that relies on correlations in general regardless of whether the features are invariant or not. Our approach uses the most confidently predicted samples identified by an OOD base model (teacher model) to train a new model (student model) that effectively adapts to the target domain. Empirical evaluations on benchmark datasets show that the performance is improved over the SOTA by ∼10-20%.