Direct importance estimation for covariate shift adaptation

Direct importance estimation for covariate shift adaptation
复制标题

DOI:
10.1007/s10463-008-0197-x
复制
发表时间:
2008-12-01
影响因子:
1
通讯作者:
Kawanabe, Motoaki
Kawanabe, Motoaki
中科院分区:
数学4区
文献类型:
--
作者:
Sugiyama, Masashi;Suzuki, Taiji;Kawanabe, Motoaki

文献摘要

被引文献

相似文献

训练样本和测试样本遵循不同输入分布的情况被称为协变量偏移。在协变量偏移情况下,诸如最大似然估计之类的标准学习方法不再一致——根据测试和训练输入密度之比的加权变体是一致的。因此,准确估计被称为重要性的密度比是协变量偏移适应中的关键问题之一。完成此任务的一种简单方法是首先分别估计训练和测试输入密度,然后通过取估计密度之比来估计重要性。然而,这种简单方法往往效果不佳,因为密度估计是一项困难的任务,特别是在高维情况下。在本文中,我们提出了一种不涉及密度估计的直接重要性估计方法。我们的方法配备了一种自然的交叉验证程序,因此可以客观地优化诸如核宽度之类的调优参数。此外,我们对所提出算法的收敛性给出了严格的数学证明。模拟实验说明了我们方法的有效性。
A situation where training and test samples follow different input distributions is called covariate shift. Under covariate shift, standard learning methods such as maximum likelihood estimation are no longer consistent-weighted variants according to the ratio of test and training input densities are consistent. Therefore, accurately estimating the density ratio, called the importance, is one of the key issues in covariate shift adaptation. A naive approach to this task is to first estimate training and test input densities separately and then estimate the importance by taking the ratio of the estimated densities. However, this naive approach tends to perform poorly since density estimation is a hard task particularly in high dimensional cases. In this paper, we propose a direct importance estimation method that does not involve density estimation. Our method is equipped with a natural cross validation procedure and hence tuning parameters such as the kernel width can be objectively optimized. Furthermore, we give rigorous mathematical proofs for the convergence of the proposed algorithm. Simulations illustrate the usefulness of our approach.