Multiple imputation of multiple multi-item scales when a full imputation model is infeasible.

Multiple imputation of multiple multi-item scales when a full imputation model is infeasible.
复制标题

DOI:
10.1186/s13104-016-1853-5
复制
发表时间:
2016-01-26
期刊:
影响因子:
1.8
通讯作者:
White IR
White IR
中科院分区:
其他
文献类型:
--
作者:
Plumpton CO;Morris T;Hughes DA;White IR

文献摘要

被引文献

相似文献

大规模调查中的数据缺失带来了重大挑战。当数据包含多个不完全多项标度时,我们重点研究用链式方程进行多重填充。最近的作者建议在单个项目的水平上对这些数据进行估算,但这可能会导致不可行的大型估算模型。我们使用从大型跨国调查收集的数据,其中分析使用九个特定国家的数据集中的每一个单独的Logistic回归模型。在这些数据中,通过链式方程对各个标度项进行多重推算在计算上是不可行的。我们提出了一种链式方程的多重归因自适应算法,通过用量表总分代替大多数量表项目来减少归因模型中的变量个数。我们评估了该方法的可行性,并将其与完整的案例分析进行了比较。我们进行了一项模拟研究,将所提出的方法与其他方法进行比较:我们在简化的设置下进行,以便与完整的补偿模型进行比较。对于实例研究,该方法将预测模型的规模从134个减少到最多72个,并使链式方程的多重推算在计算上是可行的。推算数据的分布与观测数据是一致的。多重归因的回归分析结果与完整病例分析的结果相似,但比完全病例分析的结果更精确;对于相同的回归模型,观察到标准误差减少了39%。仿真结果表明,我们提出的方法可以获得与其他方法相当的性能。通过大幅减少填充模型的规模,我们的自适应使得多重填充对于具有多个多项尺度的大比例尺调查数据是可行的。对于所考虑的数据,对多重推定数据的分析比完整的案例分析显示出更大的能力和效率。多重推算的适应更好地利用了现有的数据,可以产生与更简单的技术有很大不同的结果。本文的在线版本(doi:10.1186/s13104.0161853-5)包含补充材料,授权用户可以使用。
Missing data in a large scale survey presents major challenges. We focus on performing multiple imputation by chained equations when data contain multiple incomplete multi-item scales. Recent authors have proposed imputing such data at the level of the individual item, but this can lead to infeasibly large imputation models. We use data gathered from a large multinational survey, where analysis uses separate logistic regression models in each of nine country-specific data sets. In these data, applying multiple imputation by chained equations to the individual scale items is computationally infeasible. We propose an adaptation of multiple imputation by chained equations which imputes the individual scale items but reduces the number of variables in the imputation models by replacing most scale items with scale summary scores. We evaluate the feasibility of the proposed approach and compare it with a complete case analysis. We perform a simulation study to compare the proposed method with alternative approaches: we do this in a simplified setting to allow comparison with the full imputation model. For the case study, the proposed approach reduces the size of the prediction models from 134 predictors to a maximum of 72 and makes multiple imputation by chained equations computationally feasible. Distributions of imputed data are seen to be consistent with observed data. Results from the regression analysis with multiple imputation are similar to, but more precise than, results for complete case analysis; for the same regression models a 39 % reduction in the standard error is observed. The simulation shows that our proposed method can perform comparably against the alternatives. By substantially reducing imputation model sizes, our adaptation makes multiple imputation feasible for large scale survey data with multiple multi-item scales. For the data considered, analysis of the multiply imputed data shows greater power and efficiency than complete case analysis. The adaptation of multiple imputation makes better use of available data and can yield substantively different results from simpler techniques. The online version of this article (doi:10.1186/s13104-016-1853-5) contains supplementary material, which is available to authorized users.