Evaluation of two-fold fully conditional specification multiple imputation for longitudinal electronic health record data.

Evaluation of two-fold fully conditional specification multiple imputation for longitudinal electronic health record data.
复制标题

DOI:
10.1002/sim.6184
复制
发表时间:
2014-09-20
影响因子:
2
通讯作者:
Carpenter, James
Carpenter, James
中科院分区:
医学3区
文献类型:
--
作者:
Welch, Catherine A.;Petersen, Irene;Bartlett, Jonathan W.;White, Ian R.;Marston, Louise;Morris, Richard W.;Nazareth, Irwin;Walters, Kate;Carpenter, James

文献摘要

参考文献

被引文献

相似文献

缺失数据的多重插补 (MI) 的大多数实现都是针对简单的矩形数据结构而设计的,忽略了数据的时间顺序。因此,当将 MI 应用于具有间歇性缺失数据模式的纵向数据时,必须考虑一些替代策略。一种方法是将数据划分为时间块并在每个块独立地实现 MI。另一种方法是将所有时间块包含在同一 MI 模型中。随着时间块数量的增加,这种方法可能会因为共线性和过度拟合而失效。新的两倍完全条件规范 (FCS) MI 算法通过仅调节本地时间的测量来解决这些问题。我们描述并报告了一项新颖的模拟研究的结果,以严格评估双重 FCS 算法及其对纵向电子健康记录插补的适用性。生成完整的数据集后,大约 70% 的选定连续变量和分类变量在 10 个时间块中的每个时间块中完全随机丢失。随后,我们应用了一个简单的事件时间模型。我们比较了完整记录分析的估计系数的效率、基线时间块中的数据 MI 和两倍 FCS 算法。结果表明,两倍 FCS 算法最大限度地利用了可用数据,相对于基线 MI 的增益取决于变量内部和变量之间的相关性强度。使用这种方法还可以通过对基线值可能缺失的变量随时间进行重复测量来增加随机缺失假设的合理性。
Most implementations of multiple imputation (MI) of missing data are designed for simple rectangular data structures ignoring temporal ordering of data. Therefore, when applying MI to longitudinal data with intermittent patterns of missing data, some alternative strategies must be considered. One approach is to divide data into time blocks and implement MI independently at each block. An alternative approach is to include all time blocks in the same MI model. With increasing numbers of time blocks, this approach is likely to break down because of co-linearity and over-fitting. The new two-fold fully conditional specification (FCS) MI algorithm addresses these issues, by only conditioning on measurements, which are local in time. We describe and report the results of a novel simulation study to critically evaluate the two-fold FCS algorithm and its suitability for imputation of longitudinal electronic health records. After generating a full data set, approximately 70% of selected continuous and categorical variables were made missing completely at random in each of ten time blocks. Subsequently, we applied a simple time-to-event model. We compared efficiency of estimated coefficients from a complete records analysis, MI of data in the baseline time block and the two-fold FCS algorithm. The results show that the two-fold FCS algorithm maximises the use of data available, with the gain relative to baseline MI depending on the strength of correlations within and between variables. Using this approach also increases plausibility of the missing at random assumption by using repeated measures over time of variables whose baseline values may be missing.
流行病学和临床研究中缺少数据的多重归因:潜力和陷阱。
DOI: 10.1136/bmj.b2393
发表时间: 2009-06-29
期刊: BMJ (Clinical research ed.)
影响因子: --
作者:
Sterne JA;White IR;Carlin JB;Spratt M;Royston P;Kenward MG;Wood AM;Carpenter JR
通讯作者: Carpenter JR
DOI: 10.1371/journal.pone.0033181
发表时间: 2012-03-13
期刊: PLOS ONE
影响因子: 3.7
作者:
Wijlaars, Linda P. M. M.;Nazareth, Irwin;Petersen, Irene
通讯作者: Petersen, Irene
DOI: 10.1002/sim.3731
发表时间: 2009-12-01
影响因子: 2
作者:
Nevalainen, Jaakko;Kenward, Michael G.;Virtanen, Suvi A.
通讯作者: Virtanen, Suvi A.
DOI: 10.1080/10543401003687129
发表时间: 2011-01-01
影响因子: 1.1
作者:
Liu, G. Frank;Zhan, Xiaojiang
通讯作者: Zhan, Xiaojiang
DOI: 10.1002/mpr.330
发表时间: 2011-03
影响因子: 3.1
作者:
Grittner, Ulrike;Gmel, Gerhard;Ripatti, Samuli;Bloomfield, Kim;Wicki, Matthias
通讯作者: Wicki, Matthias