Sparse Data Reconstruction, Missing Value and Multiple Imputation through Matrix Factorization

Sparse Data Reconstruction, Missing Value and Multiple Imputation through Matrix Factorization
复制标题

DOI:
10.1177/00811750221125799
复制
发表时间:
2022-10
影响因子:
3
通讯作者:
Nandana Sengupta;Madeleine Udell;N. Srebro;James Evans
Nandana Sengupta;Madeleine Udell;N. Srebro;James Evans
中科院分区:
法学2区
文献类型:
--
作者:
Nandana Sengupta;Madeleine Udell;N. Srebro;James Evans

文献摘要

被引文献

相似文献

缺失值的社会科学方法预测从密集数据集(通常是调查)中避免的、未请求的或丢失的信息。作者提出了一种矩阵分解方法,用于缺失数据的输入(1)确定潜在因素,以模拟受访者和回应之间的相似性;(2)对各个因素进行正则化,以减少它们对最佳数据重建的过度影响。这种方法可以使社会科学家从具有大量特征的稀疏数据集中得出新的结论,例如,历史或档案来源,高流失率的在线调查,或从网络抓取创建的数据集,这些数据集与传统的归因技术相混淆。作者介绍了矩阵分解技术,并详细说明了它们的概率解释,并证明了这些技术与Rubin的多重imputation框架的一致性。作者通过使用人工数据和来自综合社会调查和全国青年纵向研究案例的真实子集的数据进行模拟,表明矩阵分解技术可能是首选。这些发现建议在几种情况下使用矩阵分解进行数据重建,特别是当数据是布尔数据和分类数据以及大部分数据缺失时。
Social science approaches to missing values predict avoided, unrequested, or lost information from dense data sets, typically surveys. The authors propose a matrix factorization approach to missing data imputation that (1) identifies underlying factors to model similarities across respondents and responses and (2) regularizes across factors to reduce their overinfluence for optimal data reconstruction. This approach may enable social scientists to draw new conclusions from sparse data sets with a large number of features, for example, historical or archival sources, online surveys with high attrition rates, or data sets created from Web scraping, which confound traditional imputation techniques. The authors introduce matrix factorization techniques and detail their probabilistic interpretation, and they demonstrate these techniques’ consistency with Rubin’s multiple imputation framework. The authors show via simulations using artificial data and data from real-world subsets of the General Social Survey and National Longitudinal Study of Youth cases for which matrix factorization techniques may be preferred. These findings recommend the use of matrix factorization for data reconstruction in several settings, particularly when data are Boolean and categorical and when large proportions of the data are missing.