Robustness of a multivariate normal approximation for imputation of incomplete binary data

Robustness of a multivariate normal approximation for imputation of incomplete binary data
复制标题

DOI:
10.1002/sim.2619
复制
发表时间:
2007-03-15
影响因子:
2
通讯作者:
Schafer, Joseph L.
Schafer, Joseph L.
中科院分区:
医学3区
文献类型:
--
作者:
Bernaards, Coen A.;Belin, Thomas R.;Schafer, Joseph L.

文献摘要

被引文献

相似文献

随着多个在多元正态模型下提供插补的软件包的出现,多重插补变得更容易执行,但缺失的二进制数据的插补仍然是一个重要的实际问题。在这里,我们探索将多元正态估算值转换为二进制估算值的三种替代方法:(1) 将估算值简单舍入到更接近 0 或 1 的值,(2) 基于“硬币翻转”的伯努利抽签,其中 0 和 I 之间的估算值被视为抽到 1 的概率,以及 (3) 自适应舍入方案,其中用于确定是否舍入到 0 或 I 的截止值基于正态近似值二项式分布,利用变量上 0 和 1 的边际比例。我们对加州健康儿童调查的 206 802 名受访者的数据集进行模拟研究,其中 198 262 人的全面观察数据定义了总体,我们从中反复抽取缺失数据的样本,进行插补,计算统计数据和置信区间,并将偏差和覆盖率与真实值进行比较。我们经常发现令人满意的偏差和覆盖特性,这表明基于统计近似的方法在应用研究中比避免发生丢失数据的情况或依赖完整案例分析更可取。考虑到覆盖范围缺陷的发生和程度,我们发现自适应舍入提供了最佳性能。版权所有 (c) 2006 John Wiley & Sons, Ltd.
Multiple imputation has become easier to perform with the advent of several software packages that provide imputations under a multivariate normal model, but imputation of missing binary data remains an important practical problem. Here, we explore three alternative methods for converting a multivariate normal imputed value into a binary imputed value: (1) simple rounding of the imputed value to the nearer of 0 or 1, (2) a Bernoulli draw based on a 'coin flip' where an imputed value between 0 and I is treated as the probability of drawing a 1, and (3) an adaptive rounding scheme where the cut-off value for determining whether to round to 0 or I is based on a normal approximation to the binomial distribution, making use of the marginal proportions of 0's and 1's on the variable. We perform simulation studies on a data set of 206 802 respondents to the California Healthy Kids Survey, where the fully observed data on 198 262 individuals defines the population, from which we repeatedly draw samples with missing data, impute, calculate statistics and confidence intervals, and compare bias and coverage against the true values. Frequently, we found satisfactory bias and coverage properties, suggesting that approaches such as these that are based on statistical approximations are preferable in applied research to either avoiding settings where missing data occur or relying on complete-case analyses. Considering both the occurrence and extent of deficits in coverage, we found that adaptive rounding provided the best performance. Copyright (c) 2006 John Wiley & Sons, Ltd.