A SEMIPARAMETRIC MULTIPLE IMPUTATION APPROACH TO FULLY SYNTHETIC DATA FOR COMPLEX SURVEYS.

A SEMIPARAMETRIC MULTIPLE IMPUTATION APPROACH TO FULLY SYNTHETIC DATA FOR COMPLEX SURVEYS.
复制标题

用于复杂调查的完全合成数据的半参数多重插补方法。

DOI:
10.1093/jssam/smac016
复制
发表时间:
2022
影响因子:
2.1
通讯作者:
Raghunathan,TrivelloreE
Raghunathan,TrivelloreE
中科院分区:
数学3区
文献类型:
--
作者:
Yu,Mandi;He,Yulei;Raghunathan,TrivelloreE

文献摘要

相似文献

数据综合是降低数据披露风险的有效统计方法。生成完全合成的数据可能会将这种风险降至最低,但对大型复杂调查的数据进行建模和应用可能会很困难。本文将两阶段的估算方法扩展到同时估算项目缺失值并生成完全合成的数据。开发了一种新的组合规则,用于使用以这种方式生成的数据进行推理。采用两种半参数缺失数据输入模型,分别对偏斜连续变量和稀疏二元变量生成完全合成的数据。采用健康与退休研究的模拟数据和真实纵向数据对所提出的方法进行了评估。并与现有的两种综合方法进行了比较:(1)采用inveware实现参数回归模型;(2)使用真实数据在synthpoppackage for R中实现的非参数分类和回归树。结果表明,使用所提出的策略,对于各种描述性和基于模型的统计数据保持了较高的数据效用。所提出的策略也比现有的复杂分析方法(如因子分析)表现得更好。
Data synthesis is an effective statistical approach for reducing data disclosure risk. Generating fully synthetic data might minimize such risk, but its modeling and application can be difficult for data from large, complex surveys. This article extended the two-stage imputation to simultaneously impute item missing values and generate fully synthetic data. A new combining rule for making inferences using data generated in this manner was developed. Two semiparametric missing data imputation models were adapted to generate fully synthetic data for skewed continuous variable and sparse binary variable, respectively. The proposed approach was evaluated using simulated data and real longitudinal data from the Health and Retirement Study. The proposed approach was also compared with two existing synthesis approaches: (1) parametric regressions models as implemented inIVEware; and (2) nonparametric Classification and Regression Trees as implemented insynthpoppackage for R using real data. The results show that high data utility is maintained for a wide variety of descriptive and model-based statistics using the proposed strategy. The proposed strategy also performs better than existing methods for sophisticated analyses such as factor analysis.