RIFLE: Imputation and Robust Inference from Low Order Marginals.

RIFLE: Imputation and Robust Inference from Low Order Marginals.
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
Transactions on machine learning research
影响因子:
--
通讯作者:
Sina Baharlouei;Kelechi Ogudu;S. Suen;Meisam Razaviyayn
Sina Baharlouei;Kelechi Ogudu;S. Suen;Meisam Razaviyayn
中科院分区:
其他
文献类型:
--
作者:
Sina Baharlouei;Kelechi Ogudu;S. Suen;Meisam Razaviyayn

文献摘要

相似文献

现实世界数据集中普遍存在的缺失值对统计推断提出了挑战,并且可以阻止在同一研究中分析类似的数据集,从而排除了许多现有数据集用于新分析的可能性。虽然已经开发了大量用于数据输入的软件包和算法,但如果存在许多缺失值和低样本量,则绝大多数都表现不佳,不幸的是,这是经验数据中的共同特征。这种低精度的估计对下游统计模型的性能有不利影响。我们开发了一个统计推理框架,用于回归和分类,在没有imputation的缺失数据的存在。我们的框架,步枪(鲁棒推断通过低阶矩估计),估计底层数据分布的低阶矩与相应的置信区间,以学习一个分布鲁棒模型。我们的框架专注于线性回归和正态判别分析,并提供收敛性和性能保证。这个框架也可以用来计算缺失的数据。在数值实验中,我们将RIFLE与几种最先进的方法(包括MICE, Amelia, MissForest, KNN-imputer, MIDA和Mean Imputer)进行比较,以在缺失值存在的情况下进行imputation和推断。我们的实验表明,当缺失值的百分比较高和/或数据点数量相对较少时,RIFLE优于其他基准算法。RIFLE可在https://github.com/optimization-for-data-driven-science/RIFLE上公开获取。
The ubiquity of missing values in real-world datasets poses a challenge for statistical inference and can prevent similar datasets from being analyzed in the same study, precluding many existing datasets from being used for new analyses. While an extensive collection of packages and algorithms have been developed for data imputation, the overwhelming majority perform poorly if there are many missing values and low sample sizes, which are unfortunately common characteristics in empirical data. Such low-accuracy estimations adversely affect the performance of downstream statistical models. We develop a statistical inference framework for regression and classification in the presence of missing data without imputation. Our framework, RIFLE (Robust InFerence via Low-order moment Estimations), estimates low-order moments of the underlying data distribution with corresponding confidence intervals to learn a distributionally robust model. We specialize our framework to linear regression and normal discriminant analysis, and we provide convergence and performance guarantees. This framework can also be adapted to impute missing data. In numerical experiments, we compare RIFLE to several state-of-the-art approaches (including MICE, Amelia, MissForest, KNN-imputer, MIDA, and Mean Imputer) for imputation and inference in the presence of missing values. Our experiments demonstrate that RIFLE outperforms other benchmark algorithms when the percentage of missing values is high and/or when the number of data points is relatively small. RIFLE is publicly available at https://github.com/optimization-for-data-driven-science/RIFLE.