Dataset Shift in Machine Learning

Dataset Shift in Machine Learning
复制标题

DOI:
10.7551/mitpress/9780262170055.001.0001
复制
发表时间:
2009-02
期刊:
影响因子:
3.9
通讯作者:
Joaquin Quionero-Candela;Masashi Sugiyama;A. Schwaighofer;Neil D. Lawrence
Joaquin Quionero-Candela;Masashi Sugiyama;A. Schwaighofer;Neil D. Lawrence
中科院分区:
环境科学与生态学3区
文献类型:
--
作者:
Joaquin Quionero-Candela;Masashi Sugiyama;A. Schwaighofer;Neil D. Lawrence

文献摘要

被引文献

相似文献

数据集偏移是预测建模中的一个常见问题,当训练和测试阶段之间输入和输出的联合分布不同时会发生这种情况。协变量移位是数据集移位的一种特殊情况,发生在只有输入分布发生变化时。数据集移位存在于大多数实际应用中,其原因从实验设计引入的偏差到训练时测试条件的不可再现性。(An例如,电子邮件垃圾邮件过滤,它可能无法识别与自动过滤器所构建的垃圾邮件形式不同的垃圾邮件。)尽管如此,尽管人们对半监督学习和主动学习的明显相似的问题给予了关注,但直到最近,数据集移位在机器学习社区中受到的关注相对较少。本卷提供了一个概述目前的努力,以处理数据集和协变量的转变。这些章节对这个问题进行了数学和哲学上的介绍,将数据集转移与迁移学习、转导、局部学习、主动学习和半监督学习相关联,提供了数据集和协变量转移的理论观点(包括决策理论和协变量转移),并提出了协变量转移的算法。参与者:Shai Ben-David,Steffen Bickel,Karsten Borgwardt,Michael Brckner,大卫Corfield,Amir Globerson,亚瑟格雷顿,Lars Kai汉森,Matthias Hein,Jiayuan Huang,Takafumi Kanamori,Klaus-Robert Mller,Sam Roweis,Neil Rubens,Tobias Scheffer,Marcel Schmittfull,Bernhard Schlkopf,Hidetoshi Shimodaira,Alex Smola,Amos Storkey,Masashi Sugiyama,Choon Hui Teo神经信息处理系列
Dataset shift is a common problem in predictive modeling that occurs when the joint distribution of inputs and outputs differs between training and test stages. Covariate shift, a particular case of dataset shift, occurs when only the input distribution changes. Dataset shift is present in most practical applications, for reasons ranging from the bias introduced by experimental design to the irreproducibility of the testing conditions at training time. (An example is -email spam filtering, which may fail to recognize spam that differs in form from the spam the automatic filter has been built on.) Despite this, and despite the attention given to the apparently similar problems of semi-supervised learning and active learning, dataset shift has received relatively little attention in the machine learning community until recently. This volume offers an overview of current efforts to deal with dataset and covariate shift. The chapters offer a mathematical and philosophical introduction to the problem, place dataset shift in relationship to transfer learning, transduction, local learning, active learning, and semi-supervised learning, provide theoretical views of dataset and covariate shift (including decision theoretic and Bayesian perspectives), and present algorithms for covariate shift. Contributors: Shai Ben-David, Steffen Bickel, Karsten Borgwardt, Michael Brckner, David Corfield, Amir Globerson, Arthur Gretton, Lars Kai Hansen, Matthias Hein, Jiayuan Huang, Takafumi Kanamori, Klaus-Robert Mller, Sam Roweis, Neil Rubens, Tobias Scheffer, Marcel Schmittfull, Bernhard Schlkopf, Hidetoshi Shimodaira, Alex Smola, Amos Storkey, Masashi Sugiyama, Choon Hui Teo Neural Information Processing series