Review and evaluation of penalised regression methods for risk prediction in low-dimensional data with few events.

Review and evaluation of penalised regression methods for risk prediction in low-dimensional data with few events.
复制标题

回顾和评估用于事件少的低维数据风险预测的惩罚回归方法。

DOI:
10.1002/sim.6782
复制
发表时间:
2016-03-30
影响因子:
2
通讯作者:
Omar RZ
Omar RZ
中科院分区:
医学3区
文献类型:
--
作者:
Pavlou M;Ambler G;Seaman S;De Iorio M;Omar RZ

文献摘要

被引文献

相似文献

风险预测模型用于使用一组预测值来预测患者的临床结果。我们专注于预测低维的二元结果,这些结果通常出现在流行病学、卫生服务和公共卫生研究中,这些领域通常使用Logistic回归。当事件的数量与回归系数的数量相比较小时,模型过度拟合可能是一个严重的问题。过度拟合的模型在应用于新数据时往往表现出较差的预测精度。我们回顾了频率收缩和贝叶斯收缩方法,它们可以通过将回归系数缩小到零来减少过拟合(一些方法也可以通过省略一些预测因素来提供更简约的模型)。我们评估了它们的预测性能,并与使用真实数据和模拟数据的最大似然估计进行了比较。模拟研究表明,在事件很少的情况下,最大似然估计往往会产生预测性能较差的过拟合模型,而惩罚方法可以提供改进。岭回归的表现很好,除了在有许多噪声预报器的情况下。Lasso在有许多噪声预测器的情况下表现得比RICE好,而在存在相关预测器的情况下表现更差。弹性网络是两者的混合体,在所有情况下都表现良好。自适应套索和平滑剪裁绝对偏差在噪声预测器较多的场景中表现最好;在其他场景中,它们的性能不如脊线和套索。当先验的超参数被仔细选择时,贝叶斯方法表现良好。它们的使用可以帮助选择变量,并且它们可以很容易地扩展到集群数据设置和合并外部信息。©2015作者。约翰威利父子有限公司出版的医学统计数据。
Risk prediction models are used to predict a clinical outcome for patients using a set of predictors. We focus on predicting low‐dimensional binary outcomes typically arising in epidemiology, health services and public health research where logistic regression is commonly used. When the number of events is small compared with the number of regression coefficients, model overfitting can be a serious problem. An overfitted model tends to demonstrate poor predictive accuracy when applied to new data. We review frequentist and Bayesian shrinkage methods that may alleviate overfitting by shrinking the regression coefficients towards zero (some methods can also provide more parsimonious models by omitting some predictors). We evaluated their predictive performance in comparison with maximum likelihood estimation using real and simulated data. The simulation study showed that maximum likelihood estimation tends to produce overfitted models with poor predictive performance in scenarios with few events, and penalised methods can offer improvement. Ridge regression performed well, except in scenarios with many noise predictors. Lasso performed better than ridge in scenarios with many noise predictors and worse in the presence of correlated predictors. Elastic net, a hybrid of the two, performed well in all scenarios. Adaptive lasso and smoothly clipped absolute deviation performed best in scenarios with many noise predictors; in other scenarios, their performance was inferior to that of ridge and lasso. Bayesian approaches performed well when the hyperparameters for the priors were chosen carefully. Their use may aid variable selection, and they can be easily extended to clustered‐data settings and to incorporate external information. © 2015 The Authors. Statistics in Medicine Published by JohnWiley & Sons Ltd.