Learning under nonstationarity: covariate shift and class‐balance change

Learning under nonstationarity: covariate shift and class‐balance change
复制标题

DOI:
10.1002/wics.1275
复制
发表时间:
2013-11
期刊:
Wiley Interdisciplinary Reviews: Computational Statistics
影响因子:
--
通讯作者:
Masashi Sugiyama;M. Yamada;M. C. D. Plessis
Masashi Sugiyama;M. Yamada;M. C. D. Plessis
中科院分区:
其他
文献类型:
--
作者:
Masashi Sugiyama;M. Yamada;M. C. D. Plessis

文献摘要

相似文献

许多监督式机器学习算法背后的基本假设之一是训练和测试数据遵循相同的概率分布。然而,这一重要的假设在实践中经常被违反,例如,由于不可避免的样本选择偏差或环境的非平稳性。由于违反了这一假设,标准的机器学习方法会出现显著的估计偏差。在本文中,我们考虑了这种分布变化的两种情况-输入分布不同的协变量变化和分类先验概率在分类中变化的类平衡变化-并回顾了基于重要性加权的半监督自适应技术。WIREs Comput Stat 2013,5:465-477。doi:10.1002/wics.1275
One of the fundamental assumptions behind many supervised machine‐learning algorithms is that training and test data follow the same probability distribution. However, this important assumption is often violated in practice, for example, because of an unavoidable sample selection bias or nonstationarity of the environment. Owing to violation of the assumption, standard machine‐learning methods suffer a significant estimation bias. In this article, we consider two scenarios of such distribution change—the covariate shift where input distributions differ and class‐balance change where class‐prior probabilities vary in classification—and review semi‐supervised adaptation techniques based on importance weighting. WIREs Comput Stat 2013, 5:465–477. doi: 10.1002/wics.1275