Evaluation of two-fold fully conditional specification multiple imputation for longitudinal electronic health record data.
Evaluation of two-fold fully conditional specification multiple imputation for longitudinal electronic health record data.
复制标题
DOI:
10.1002/sim.6184
复制
发表时间:
2014-09-20
影响因子:
2
通讯作者:
Carpenter, James
中科院分区:
文献类型:
--
作者:
Welch, Catherine A.;Petersen, Irene;Bartlett, Jonathan W.;White, Ian R.;Marston, Louise;Morris, Richard W.;Nazareth, Irwin;Walters, Kate;Carpenter, James
Most implementations of multiple imputation (MI) of missing data are designed for simple rectangular data structures ignoring temporal ordering of data. Therefore, when applying MI to longitudinal data with intermittent patterns of missing data, some alternative strategies must be considered. One approach is to divide data into time blocks and implement MI independently at each block. An alternative approach is to include all time blocks in the same MI model. With increasing numbers of time blocks, this approach is likely to break down because of co-linearity and over-fitting. The new two-fold fully conditional specification (FCS) MI algorithm addresses these issues, by only conditioning on measurements, which are local in time. We describe and report the results of a novel simulation study to critically evaluate the two-fold FCS algorithm and its suitability for imputation of longitudinal electronic health records. After generating a full data set, approximately 70% of selected continuous and categorical variables were made missing completely at random in each of ten time blocks. Subsequently, we applied a simple time-to-event model. We compared efficiency of estimated coefficients from a complete records analysis, MI of data in the baseline time block and the two-fold FCS algorithm. The results show that the two-fold FCS algorithm maximises the use of data available, with the gain relative to baseline MI depending on the strength of correlations within and between variables. Using this approach also increases plausibility of the missing at random assumption by using repeated measures over time of variables whose baseline values may be missing.
登录
查看更多内容
DOI:
10.1136/bmj.b2393
发表时间:
2009-06-29
期刊:
BMJ (Clinical research ed.)
影响因子:
--
作者:
Sterne JA;White IR;Carlin JB;Spratt M;Royston P;Kenward MG;Wood AM;Carpenter JR
通讯作者:
Carpenter JR
影响因子:
3.7
作者:
Wijlaars, Linda P. M. M.;Nazareth, Irwin;Petersen, Irene
通讯作者:
Petersen, Irene
影响因子:
2
作者:
Nevalainen, Jaakko;Kenward, Michael G.;Virtanen, Suvi A.
通讯作者:
Virtanen, Suvi A.
影响因子:
1.1
作者:
Liu, G. Frank;Zhan, Xiaojiang
通讯作者:
Zhan, Xiaojiang
影响因子:
3.1
作者:
Grittner, Ulrike;Gmel, Gerhard;Ripatti, Samuli;Bloomfield, Kim;Wicki, Matthias
通讯作者:
Wicki, Matthias