SPECIAL SERIES: MISSING DATA Review: A gentle introduction to imputation of missing values

SPECIAL SERIES: MISSING DATA Review: A gentle introduction to imputation of missing values
复制标题

DOI:
--
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
A. Rogier;T. Donders;T. Stijnen
A. Rogier;T. Donders;T. Stijnen
中科院分区:
其他
文献类型:
--
作者:
A. Rogier;T. Donders;T. Stijnen

文献摘要

被引文献

相似文献

在大多数情况下,处理缺失数据的简单技术(如完整案例分析、总体平均值推算和缺失指标法)会产生有偏见的结果,而推算技术在进行推算后会产生有效的结果,而不会使分析复杂化。归罪技术基于这样一种想法,即研究样本中的任何对象都可以被从相同来源人群中随机选择的新对象取代。对变量的缺失数据的归因是用从该变量的分布估计中提取的值来替换缺失的数据。在单一推算中,只使用一个估计值。在多重补偿中,使用了不同的估计,反映了该分布估计的不确定性。在所谓的随机缺失和完全随机缺失的一般情况下,无论是单一的还是多重的归因都会导致对研究联系的无偏估计。但是,单次推算导致的估计标准误差太小,而多次推算导致正确估计的标准误差和可信区间。在本文中,我们将解释为什么会出现这种情况,并使用一个简单的模拟研究来演示我们的解释。我们还解释和说明了为什么两种常用的处理缺失数据的方法,即总体均值推算和
In most situations, simple techniques for handling missing data (such as complete case analysis, overall mean imputation, and the missing-indicator method) produce biased results, whereas imputation techniques yield valid results without complicating the analysis once the imputations are carried out. Imputation techniques are based on the idea that any subject in a study sample can be replaced by a new randomly chosen subject from the same source population. Imputation of missing data on a variable is replacing that missing by a value that is drawn from an estimate of the distribution of this variable. In single imputation, only one estimate is used. In multiple imputation, various estimates are used, reflecting the uncertainty in the estimation of this distribution. Under the general conditions of so-called missing at random and missing completely at random, both single and multiple imputations result in unbiased estimates of study associations. But single imputation results in too small estimated standard errors, whereas multiple imputation results in correctly estimated standard errors and confidence intervals. In this article we explain why all this is the case, and use a simple simulation study to demonstrate our explanations. We also explain and illustrate why two frequently used methods to handle missing data, i.e., overall mean imputation and