Low-Dimensional Density Ratio Estimation for Covariate Shift Correction

Low-Dimensional Density Ratio Estimation for Covariate Shift Correction
复制标题

DOI:
--
复制
发表时间:
2019-04
期刊:
Proceedings of machine learning research
影响因子:
--
通讯作者:
P. Stojanov;Mingming Gong;J. Carbonell;Kun Zhang
P. Stojanov;Mingming Gong;J. Carbonell;Kun Zhang
中科院分区:
其他
文献类型:
--
作者:
P. Stojanov;Mingming Gong;J. Carbonell;Kun Zhang

文献摘要

被引文献

相似文献

协变量移位是监督学习的一种普遍设置,当训练和测试数据来自不同的时间段、不同但相关的领域或通过不同的采样策略时。本文讨论了迁移学习设置,与源和目标域之间的协变量转移。现有的协变量偏移校正方法大多利用特征的密度比来对源域数据进行重新加权,当特征为高维时,估计的密度比可能存在较大的估计方差,导致预测性能不佳.在这项工作中,我们调查的协变量移位校正性能的特征的维数的依赖性,并提出了一种校正方法,找到一个低维的特征表示,考虑到相关的目标Y的功能,并利用密度比的重要性重新加权表示。我们讨论了影响我们的方法的性能的因素,并展示了它的能力,伪真实和真实世界的数据。
Covariate shift is a prevalent setting for supervised learning in the wild when the training and test data are drawn from different time periods, different but related domains, or via different sampling strategies. This paper addresses a transfer learning setting, with covariate shift between source and target domains. Most existing methods for correcting covariate shift exploit density ratios of the features to reweight the source-domain data, and when the features are high-dimensional, the estimated density ratios may suffer large estimation variances, leading to poor prediction performance. In this work, we investigate the dependence of covariate shift correction performance on the dimensionality of the features, and propose a correction method that finds a low-dimensional representation of the features, which takes into account feature relevant to the target Y, and exploits the density ratio of this representation for importance reweighting. We discuss the factors affecting the performance of our method and demonstrate its capabilities on both pseudo-real and real-world data.