On the consistency of supervised learning with missing values

On the consistency of supervised learning with missing values
复制标题

DOI:
--
复制
发表时间:
2019-02
期刊:
ArXiv
影响因子:
--
通讯作者:
J. Josse;Nicolas Prost;Erwan Scornet;G. Varoquaux
J. Josse;Nicolas Prost;Erwan Scornet;G. Varoquaux
中科院分区:
其他
文献类型:
--
作者:
J. Josse;Nicolas Prost;Erwan Scornet;G. Varoquaux

文献摘要

被引文献

相似文献

在许多应用程序设置中,数据有缺失的条目,这使得分析具有挑战性。大量的文献在推理框架中解决了缺失值:从不完整的表中估计参数及其方差。在这里,我们考虑监督学习设置:当训练和测试数据中都出现缺失值时预测目标。我们证明了两种方法在预测中的一致性。一个引人注目的结果是,当缺失值不具有信息性时,广泛使用的用常量(例如学习前的平均值)进行输入的方法是一致的。这与推断设置形成对比,其中平均imputation被指向扭曲数据分布。这种简单的方法在实践中可以保持一致,这一点很重要。我们还表明,适合于完整观测的预测器可以通过多次imputation对不完整数据进行最佳预测。最后,为了比较直接学习的imputation与解释缺失值的模型,我们进一步分析了决策树。由于它们能够处理不完全变量的半离散性质,它们可以自然地解决带有缺失值的经验风险最小化问题。在比较了理论和经验上不同的树缺失值策略后,我们推荐使用“属性缺失”方法,因为它既可以处理非信息性缺失值,也可以处理信息性缺失值。
In many application settings, the data have missing entries which make analysis challenging. An abundant literature addresses missing values in an inferential framework: estimating parameters and their variance from incomplete tables. Here, we consider supervised-learning settings: predicting a target when missing values appear in both training and testing data. We show the consistency of two approaches in prediction. A striking result is that the widely-used method of imputing with a constant, such as the mean prior to learning is consistent when missing values are not informative. This contrasts with inferential settings where mean imputation is pointed at for distorting the distribution of the data. That such a simple approach can be consistent is important in practice. We also show that a predictor suited for complete observations can predict optimally on incomplete data, through multiple imputation. Finally, to compare imputation with learning directly with a model that accounts for missing values, we analyze further decision trees. These can naturally tackle empirical risk minimization with missing values, due to their ability to handle the half-discrete nature of incomplete variables. After comparing theoretically and empirically different missing values strategies in trees, we recommend using the "missing incorporated in attribute" method as it can handle both non-informative and informative missing values.