Learning from incomplete data with infinite imputations

Learning from incomplete data with infinite imputations
复制标题

DOI:
10.1145/1390156.1390186
复制
发表时间:
2008-07
期刊:
--
影响因子:
--
通讯作者:
Uwe Dick;P. Haider;T. Scheffer
Uwe Dick;P. Haider;T. Scheffer
中科院分区:
其他
文献类型:
--
作者:
Uwe Dick;P. Haider;T. Scheffer

文献摘要

被引文献

相似文献

我们解决了从训练数据中学习决策函数的问题,其中一些属性值是不可观察的。例如,当训练数据从多个源聚合时,并且某些源仅记录属性的子集时,可能会出现此问题。我们推导出一个通用的联合优化问题,其中缺失值的分布是一个自由参数。我们表明,最优解集中的密度质量上的许多插补,并提供了相应的算法,从不完整的数据学习。我们报告的实证结果基准数据,并在电子邮件垃圾邮件的应用程序,激励我们的工作。
We address the problem of learning decision functions from training data in which some attribute values are unobserved. This problem can arise, for instance, when training data is aggregated from multiple sources, and some sources record only a subset of attributes. We derive a generic joint optimization problem in which the distribution governing the missing values is a free parameter. We show that the optimal solution concentrates the density mass on finitely many imputations, and provide a corresponding algorithm for learning from incomplete data. We report on empirical results on benchmark data, and on the email spam application that motivates our work.