Events per variable (EPV) and the relative performance of different strategies for estimating the out-of-sample validity of logistic regression models.

Events per variable (EPV) and the relative performance of different strategies for estimating the out-of-sample validity of logistic regression models.
复制标题

DOI:
10.1177/0962280214558972
复制
发表时间:
2017-04
影响因子:
2.3
通讯作者:
Steyerberg EW
Steyerberg EW
中科院分区:
医学3区
文献类型:
--
作者:
Austin PC;Steyerberg EW

文献摘要

被引文献

相似文献

我们进行了一系列广泛的实证分析,以研究每个变量的事件数(EPV)对三种不同方法的相对性能的影响,这些方法用于评估逻辑回归模型的预测准确性:分析样本中的表观性能,分裂样本验证和使用自举方法的乐观校正。使用一个单一的数据集的心力衰竭住院患者,我们比较了这些方法的歧视性能的估计值,从同一人群中产生的一个非常大的独立验证样本。正如预期的那样,表观表现存在乐观偏差,随着每个变量的事件数量增加,乐观程度逐渐降低。一旦每个变量的事件数至少为20,则自举校正方法和使用独立验证样本之间的差异最小。分裂样本的评估导致过于悲观和高度不确定的模型性能的估计。与分裂样本估计相比,表观性能估计的均方误差较低,但最低的均方误差是通过自举校正乐观估计获得的。对于性能估计值的偏倚、方差和均方误差,使用分裂样本验证产生的惩罚相当于按比例减少样本量,该比例相当于保留用于模型验证的样本比例。总之,分裂样本验证是低效的,表观表现过于乐观的内部验证基于回归的预测模型。现代验证方法,如基于bootstrap的乐观校正,是可取的。虽然这些发现对许多统计学家来说可能并不奇怪,但当前研究的结果加强了在临床预测模型的开发和验证中应该被认为是良好的统计实践。
We conducted an extensive set of empirical analyses to examine the effect of the number of events per variable (EPV) on the relative performance of three different methods for assessing the predictive accuracy of a logistic regression model: apparent performance in the analysis sample, split-sample validation, and optimism correction using bootstrap methods. Using a single dataset of patients hospitalized with heart failure, we compared the estimates of discriminatory performance from these methods to those for a very large independent validation sample arising from the same population. As anticipated, the apparent performance was optimistically biased, with the degree of optimism diminishing as the number of events per variable increased. Differences between the bootstrap-corrected approach and the use of an independent validation sample were minimal once the number of events per variable was at least 20. Split-sample assessment resulted in too pessimistic and highly uncertain estimates of model performance. Apparent performance estimates had lower mean squared error compared to split-sample estimates, but the lowest mean squared error was obtained by bootstrap-corrected optimism estimates. For bias, variance, and mean squared error of the performance estimates, the penalty incurred by using split-sample validation was equivalent to reducing the sample size by a proportion equivalent to the proportion of the sample that was withheld for model validation. In conclusion, split-sample validation is inefficient and apparent performance is too optimistic for internal validation of regression-based prediction models. Modern validation methods, such as bootstrap-based optimism correction, are preferable. While these findings may be unsurprising to many statisticians, the results of the current study reinforce what should be considered good statistical practice in the development and validation of clinical prediction models.