Synthetic Multiple-Imputation Procedure for Multistage Complex Samples.

Synthetic Multiple-Imputation Procedure for Multistage Complex Samples.
复制标题

DOI:
10.1515/jos-2016-0011
复制
发表时间:
2016-03
影响因子:
1.1
通讯作者:
Raghunathan TE
Raghunathan TE
中科院分区:
数学4区
文献类型:
--
作者:
Zhou H;Elliott MR;Raghunathan TE

文献摘要

相似文献

当存在项目级缺失数据时,通常使用多重插补(MI)。然而,多元智能要求将调查设计信息纳入估算模型。对于多阶段分层聚类设计,这需要虚拟变量来代表层以及插补模型中每个层内嵌套的主要抽样单位(PSU)。这样的建模策略不仅操作繁琐,而且在样本设计中有许多层时效率低下。复杂性仅在需要对采样权重建模时才增加。本文提出了一种通用的分析策略,用于从具有项目级缺失的复杂样本设计中进行总体推断。在模拟研究中,所提出的程序表现出有效的估计和良好的覆盖性能。我们还考虑了一个应用程序,以适应缺失的身体质量指数(BMI)数据的分析中使用的国家健康和营养检查调查(NHANES)III数据的BMI指数。我们认为,所提出的方法提供了一个易于实现的解决方案,目前的MI技术没有很好地处理的问题。请注意,虽然所提出的方法借用MI框架来开发其推理方法,但它并不是作为一种替代策略来发布复杂样本设计数据的多重插补数据集,而是作为一种分析策略本身。
Multiple imputation (MI) is commonly used when item-level missing data are present. However, MI requires that survey design information be built into the imputation models. For multistage stratified clustered designs, this requires dummy variables to represent strata as well as primary sampling units (PSUs) nested within each stratum in the imputation model. Such a modeling strategy is not only operationally burdensome but also inferentially inefficient when there are many strata in the sample design. Complexity only increases when sampling weights need to be modeled. This article develops a general-purpose analytic strategy for population inference from complex sample designs with item-level missingness. In a simulation study, the proposed procedures demonstrate efficient estimation and good coverage properties. We also consider an application to accommodate missing body mass index (BMI) data in the analysis of BMI percentiles using National Health and Nutrition Examination Survey (NHANES) III data. We argue that the proposed methods offer an easy-to-implement solution to problems that are not well-handled by current MI techniques. Note that, while the proposed method borrows from the MI framework to develop its inferential methods, it is not designed as an alternative strategy to release multiply imputed datasets for complex sample design data, but rather as an analytic strategy in and of itself.