Effect of Variable Selection Strategy on the Performance of Prognostic Models When Using Multiple Imputation.

Effect of Variable Selection Strategy on the Performance of Prognostic Models When Using Multiple Imputation.
复制标题

DOI:
10.1161/circoutcomes.119.005927
复制
发表时间:
2019-11
期刊:
Circulation. Cardiovascular quality and outcomes
影响因子:
--
通讯作者:
White IR
White IR
中科院分区:
其他
文献类型:
--
作者:
Austin PC;Lee DS;Ko DT;White IR

文献摘要

被引文献

相似文献

变量选择是建立预后模型时的一个重要问题。临床研究中经常会出现数据缺失的情况。多重归因越来越多地被用来解决临床研究中缺失数据的存在。不同的变量选择策略与多重输入数据对衍生预测模型的外部性能的影响还没有得到很好的检验。我们使用9种不同方法的反向变量选择来处理派生样本中的多重输入数据,以建立预测急性心肌梗死患者住院1年内死亡的Logistic回归模型。我们在时间上不同的验证样本中评估了每个衍生模型的预测准确性。衍生和验证样本分别包括1999年至2001年间住院的11名 524名患者和2004年至2005年间住院的7889名患者。我们考虑了41个候选预测变量。数据缺失的情况经常发生,派生样本中只有13%的患者和验证样本中的31%的患者拥有完整的数据。不管变量选择的显著程度如何,仅使用派生样本中的完整病例开发的预后模型在验证样本中的表现比使用派生样本的多个推定版本为其选择变量的模型的表现要差得多。其他8种处理多重推论数据的方法得到的预后模型的性能彼此相似。忽略缺失的数据,只使用具有完整数据的受试者,可能会导致预测模型的推导性能较差。在建立预后模型时,应使用多重补偿来解释丢失的数据。
Variable selection is an important issue when developing prognostic models. Missing data occur frequently in clinical research. Multiple imputation is increasingly used to address the presence of missing data in clinical research. The effect of different variable selection strategies with multiply imputed data on the external performance of derived prognostic models has not been well examined. We used backward variable selection with 9 different ways to handle multiply imputed data in a derivation sample to develop logistic regression models for predicting death within 1 year of hospitalization with an acute myocardial infarction. We assessed the prognostic accuracy of each derived model in a temporally distinct validation sample. The derivation and validation samples consisted of 11 524 patients hospitalized between 1999 and 2001 and 7889 patients hospitalized between 2004 and 2005, respectively. We considered 41 candidate predictor variables. Missing data occurred frequently, with only 13% of patients in the derivation sample and 31% of patients in the validation sample having complete data. Regardless of the significance level for variable selection, the prognostic model developed using only the complete cases in the derivation sample had substantially worse performance in the validation sample than did the models for which variables were selected using the multiply imputed versions of the derivation sample. The other 8 approaches to handling multiply imputed data resulted in prognostic models with performance similar to one another. Ignoring missing data and using only subjects with complete data can result in the derivation of prognostic models with poor performance. Multiple imputation should be used to account for missing data when developing prognostic models.