Integration of data from probability surveys and big found data for finite population inference using mass imputation

Integration of data from probability surveys and big found data for finite population inference using mass imputation
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Jae Kwang Kim;Y. Hwang;Paul H. Chook;Shu Yang
Jae Kwang Kim;Y. Hwang;Paul H. Chook;Shu Yang
中科院分区:
其他
文献类型:
--
作者:
Jae Kwang Kim;Y. Hwang;Paul H. Chook;Shu Yang

文献摘要

被引文献

相似文献

在大数据时代,越来越多的数据来源可以用于统计分析。作为有限总体推理中的一个重要例子,我们考虑了一种将概率调查数据与大发现数据相结合的归算方法。我们关注的是研究变量只在大数据中被观察到,而其他辅助变量在两个数据中都被观察到的情况。与通常的缺失数据分析的输入不同,我们为概率样本中的所有单位创建输入值。在调查数据整合的背景下,这种大规模的imputation是有吸引力的(Kim和Rao, 2012)。我们将大规模归算扩展为调查数据与非调查大数据的数据整合工具。介绍了质量归算方法及其统计性质。Rivers(2007)的匹配估计量也作为一个特例进行了讨论。讨论了质量输入数据的方差估计。仿真结果表明,所提出的估计器在鲁棒性和效率方面都优于现有的竞争对手。
Multiple data sources are becoming increasingly available for statistical analyses in the era of big data. As an important example in finite-population inference, we consider an imputation approach to combining data from a probability survey and big found data. We focus on the case when the study variable is observed in the big data only, but the other auxiliary variables are commonly observed in both data. Unlike the usual imputation for missing data analysis, we create imputed values for all units in the probability sample. Such mass imputation is attractive in the context of survey data integration (Kim and Rao, 2012). We extend mass imputation as a tool for data integration of survey data and big non-survey data. The mass imputation methods and their statistical properties are presented. The matching estimator of Rivers (2007) is also covered as a special case. Variance estimation with mass-imputed data is discussed. The simulation results demonstrate the proposed estimators outperform existing competitors in terms of robustness and efficiency.