Learning from incomplete data with infinite imputations
Learning from incomplete data with infinite imputations
复制标题
DOI:
10.1145/1390156.1390186
复制
发表时间:
2008-07
期刊:
影响因子:
--
通讯作者:
Uwe Dick;P. Haider;T. Scheffer
中科院分区:
文献类型:
--
作者:
Uwe Dick;P. Haider;T. Scheffer
We address the problem of learning decision functions from training data in which some attribute values are unobserved. This problem can arise, for instance, when training data is aggregated from multiple sources, and some sources record only a subset of attributes. We derive a generic joint optimization problem in which the distribution governing the missing values is a free parameter. We show that the optimal solution concentrates the density mass on finitely many imputations, and provide a corresponding algorithm for learning from incomplete data. We report on empirical results on benchmark data, and on the email spam application that motivates our work.