A potential for bias when rounding in multiple imputation

A potential for bias when rounding in multiple imputation
复制标题

DOI:
10.1198/0003130032314
复制
发表时间:
2003-11-01
影响因子:
1.8
通讯作者:
Parzen, M
Parzen, M
中科院分区:
数学2区
文献类型:
--
作者:
Horton, NJ;Lipsitz, SR;Parzen, M

文献摘要

被引文献

相似文献

随着支持多重插补以分析缺失数据的数据集的通用软件包(例如 Solas、SAS PROC MI. 和 S-Plus 6.0)的出现,我们预计未来会更多地使用多重插补。为简单起见,一些插补包假设多重插补模型中变量的联合分布是多元正态分布,并根据给定观测数据的缺失数据的条件正态分布插补缺失数据。如果可能丢失的数据不是多元正态数据(例如二元数据),则输入正态随机变量可能会产生难以置信的值。为了解决这个问题,人们开发了多种方法,包括将估算法线四舍五入到数据集中最接近的观测值。我们表明,这种舍入可能会导致参数估计出现偏差,而如果估算值不进行舍入,则不会发生偏差。本文表明,不应随意使用舍入,因此在舍入估算值时应谨慎行事,特别是对于二分变量。
With the advent of general purpose packages that support multiple imputation for analyzing datasets with missing data (e.g., Solas, SAS PROC MI. and S-Plus 6.0), we expect much greater use of multiple imputation in the future. For simplicity, some imputation packages assume the joint distribution of the variables in the multiple imputation model is multivariate normal, and impute the missing data from the conditional normal distribution for the missing data given the observed data. If the possibly missing data are not multivariate normal (say, binary), imputing a normal random variable can yield implausible values. To circumvent this problem, a number of methods have been developed, including rounding the imputed normal to the closest observed value in the dataset. We show that this rounding can cause biased estimates of parameters, whereas if the imputed value is not rounded, no bias would occur. This article shows that rounding should not be used indiscriminately, and thus some caution should be exercised when rounding imputed values, particularly for dichotomous variables.