Near-optimal rate of consistency for linear models with missing values

Near-optimal rate of consistency for linear models with missing values
复制标题

具有缺失值的线性模型的接近最佳一致性率

DOI:
--
复制
发表时间:
2022
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
Erwan Scornet
Erwan Scornet
中科院分区:
--
文献类型:
--
作者:
Alexis Ayme;Claire Boyer;Aymeric Dieuleveut;Erwan Scornet

文献摘要

被引文献

相似文献

在大多数真实世界的数据集中,由于多个来源的聚集和本质上缺失的信息(传感器故障,调查中未回答的问题……),会出现缺失值。事实上,缺失值的本质通常会阻止我们运行标准的学习算法。在本文中,我们关注的是被广泛研究的线性模型,但在存在缺失值的情况下,这是一项相当具有挑战性的任务。实际上,贝叶斯预测器可以分解为每个缺失模式对应的预测器之和。这最终需要解决大量的学习任务,输入特征的数量呈指数增长,这使得对当前现实世界数据集的预测变得不可能。首先,我们提出了一个严格的设定来分析最小二乘型估计量,并建立了在维数上呈指数增长的超额风险的界。因此,我们利用缺失的数据分布提出了一种新的算法,并推导出相关的自适应风险界限,结果证明是最小最大最优的。数值实验突出了与用于缺失值预测的最先进算法相比,我们的方法的优点。
Missing values arise in most real-world data sets due to the aggregation of multiple sources and intrinsically missing information (sensor failure, unanswered questions in surveys...). In fact, the very nature of missing values usually prevents us from running standard learning algorithms. In this paper, we focus on the extensively-studied linear models, but in presence of missing values, which turns out to be quite a challenging task. In-deed, the Bayes predictor can be decomposed as a sum of predictors corresponding to each missing pattern. This eventually requires to solve a number of learning tasks, exponential in the number of input features, which makes predictions impossible for current real-world datasets. First, we propose a rigorous setting to analyze a least-square type estimator and establish a bound on the excess risk which increases exponentially in the dimension. Consequently, we leverage the missing data distribution to propose a new algo-rithm, and derive associated adaptive risk bounds that turn out to be minimax optimal. Numerical experiments highlight the benefits of our method compared to state-of-the-art algorithms used for predictions with missing values.