External validation of prognostic models for critically ill patients required substantial sample sizes

External validation of prognostic models for critically ill patients required substantial sample sizes
复制标题

DOI:
10.1016/j.jclinepi.2006.08.011
复制
发表时间:
2007-05-01
影响因子:
7.2
通讯作者:
de Keizer, N. F.
de Keizer, N. F.
中科院分区:
医学2区
文献类型:
--
作者:
Peek, N.;Arts, D. G. T.;de Keizer, N. F.

文献摘要

被引文献

相似文献

目的:研究在重症监护病房(icu)预后模型外部验证中常用的预测性能指标的行为。研究设计和背景:四种预后模型(简化急性生理评分II、急性生理和慢性健康评估II和死亡概率模型II)在荷兰国家重症监护评估注册数据库中进行评估。对于每个模型判别(AUC),准确性(Brier评分)和两个校准措施,对来自41,239名ICU入院患者的数据进行评估。对从数据库中随机抽取的较小的子样本重复此验证过程,并将结果与在整个数据集上获得的结果进行比较。结果:模型间的性能差异较小。AUC和Brier评分在小样本中变化较大。AUC值的标准误差准确,但检测性能差异的能力较低。校准试验对样本量极为敏感。在没有统计分析的情况下,两种方法的直接比较都是不可靠的。结论:外部验证的性能评估和模型比较需要大量的样本量。不应在这些设置中使用校准统计和显著性检验。相反,建议使用一种简单的定制方法来修复不匹配问题。(c) 2007爱思唯尔公司版权所有。
Objective: To investigate the behavior of predictive performance measures that are commonly used in external validation of prognostic models for outcome at intensive care units (ICUs).Study Design and Setting: Four prognostic models (Simplified Acute Physiology Score II, the Acute Physiology and Chronic Health Evaluation II, and the Mortality Probability Models II) were evaluated in the Dutch National Intensive Care Evaluation registry database. For each model discrimination (AUC), accuracy (Brier score), and two calibration measures were assessed on data from 41,239 ICU admissions. This validation procedure was repeated with smaller subsamples randomly drawn from the database, and the results were compared with those obtained on the entire data set.Results: Differences in performance between the models were small. The AUC and Brier score showed large variation with small samples. Standard errors of AUC values were accurate but the power to detect differences in performance was low. Calibration tests were extremely sensitive to sample size. Direct comparison of performance, without statistical analysis, was unreliable with either measure.Conclusion: Substantial sample sizes are required for performance assessment and model comparison in external validation. Calibration statistics and significance tests should not be used in these settings. Instead, a simple custornization method to repair lack-of-fit problems is recommended. (c) 2007 Elsevier Inc. All rights reserved.