Missing Data in Clinical Research: A Tutorial on Multiple Imputation.

Missing Data in Clinical Research: A Tutorial on Multiple Imputation.
复制标题

临床研究中的缺失数据:多重插补教程。

DOI:
10.1016/j.cjca.2020.11.010
复制
发表时间:
2021-09
期刊:
The Canadian journal of cardiology
影响因子:
--
通讯作者:
van Buuren S
van Buuren S
中科院分区:
其他
文献类型:
--
作者:
Austin PC;White IR;Lee DS;van Buuren S

文献摘要

参考文献

被引文献

相似文献

缺失数据是临床研究中常见的现象。缺失数据发生在未测量或记录样本中所有受试者的感兴趣变量值时。解决缺失数据的常见方法包括完整病例分析(排除缺失数据的受试者)和均值插补(用未缺失的受试者中该变量的均值替换缺失值)。然而,在许多情况下,这些方法可能导致统计数据(例如回归系数)和/或人为狭窄的置信区间的有偏估计。多重插补(MI)是解决缺失数据存在的一种流行方法。对于MI,对于每例缺失该变量数据的受试者,插补或填写给定变量的多个合理值。这导致创建多个完整的数据集。在这些完整数据集中的每一个中进行相同的统计分析,并在完整数据集中汇总结果。我们提供了一个介绍MI和讨论的问题,在其实施中,包括开发插补模型,有多少插补数据集创建,并解决派生变量。我们通过对心力衰竭住院患者数据的分析来说明MI的应用。我们专注于开发一个模型,以估计在缺失数据的情况下1年死亡率的概率。提供了在R、SAS和Stata中进行MI的统计软件代码。
Missing data is a common occurrence in clinical research. Missing data occurs when the value of the variables of interest are not measured or recorded for all subjects in the sample. Common approaches to addressing the presence of missing data include complete-case analyses, where subjects with missing data are excluded, and mean-value imputation, where missing values are replaced with the mean value of that variable in those subjects for whom it is not missing. However, in many settings, these approaches can lead to biased estimates of statistics (eg, of regression coefficients) and/or confidence intervals that are artificially narrow. Multiple imputation (MI) is a popular approach for addressing the presence of missing data. With MI, multiple plausible values of a given variable are imputed or filled in for each subject who has missing data for that variable. This results in the creation of multiple completed data sets. Identical statistical analyses are conducted in each of these complete data sets and the results are pooled across complete data sets. We provide an introduction to MI and discuss issues in its implementation, including developing the imputation model, how many imputed data sets to create, and addressing derived variables. We illustrate the application of MI through an analysis of data on patients hospitalised with heart failure. We focus on developing a model to estimate the probability of 1-year mortality in the presence of missing data. Statistical software code for conducting MI in R, SAS, and Stata are provided.
流行病学和临床研究中缺少数据的多重归因:潜力和陷阱。
DOI: 10.1136/bmj.b2393
发表时间: 2009-06-29
期刊: BMJ (Clinical research ed.)
影响因子: --
作者:
Sterne JA;White IR;Carlin JB;Spratt M;Royston P;Kenward MG;Wood AM;Carpenter JR
通讯作者: Carpenter JR
DOI: 10.1186/1471-2288-12-46
发表时间: 2012-04-10
影响因子: 4
作者:
Seaman SR;Bartlett JW;White IR
通讯作者: White IR
DOI: 10.1037//1082-989x.7.2.147
发表时间: 2002-06-01
影响因子: 7
作者:
Schafer, JL;Graham, JW
通讯作者: Graham, JW
DOI: 10.1002/sim.1981
发表时间: 2005-04-15
影响因子: 2
作者:
White, IR;Thompson, SG
通讯作者: Thompson, SG
DOI: 10.1002/sim.3618
发表时间: 2009-07-10
影响因子: 2
作者:
White, Ian R.;Royston, Patrick
通讯作者: Royston, Patrick