Prognostic modeling with logistic regression analysis: In search of a sensible strategy in small data sets

Prognostic modeling with logistic regression analysis: In search of a sensible strategy in small data sets
复制标题

DOI:
10.1177/0272989x0102100106
复制
发表时间:
2001-01-01
影响因子:
3.6
通讯作者:
Habbema, JDF
Habbema, JDF
中科院分区:
医学3区
文献类型:
--
作者:
Steyerberg, EW;Eijkemans, MJC;Habbema, JDF

文献摘要

被引文献

相似文献

临床决策通常需要估计单个专利中二分结果的可能性。当有经验数据时,这些估计值很可能从逻辑回归模型中获得。在发展这种模式时可以采取几种战略。在这项研究中,作者比较了来自急性心肌梗死患者大数据集的23个小样本的替代策略,他们开发了30天死亡率的预测模型。在数据集的独立部分进行评价。具体而言,作者研究了协变量的编码和逐步选择对所得模型的判别能力的影响,以及统计“收缩”技术对校准的影响。正如预期的那样,连续协变量的二分法意味着信息的丢失。值得注意的是,与包括所有可用协变量的完整模型相比,逐步选择导致模型的区分度较低,即使其中一半以上与结果随机相关。使用定性信息对预测因子的预测能力略有提高。当收缩应用于回归系数的标准最大似然估计时,校准得到改善。总之,在小数据集中,一个明智的策略是在完整的模型中应用收缩方法,其中包括基于外部信息选择的良好编码的预测因子。
Clinical decision making often requires estimates of the likelihood of a dichotomous outcome in individual patents. When empirical data are available, these estimates may well be obtained from a logistic regression model. Several strategies may be followed in the development of such a model. In this study, the authors compare alternative strategies in 23 small subsamples from a large data set of patients with an acute myocardial infarction, where they developed predictive models for 30-day mortality. Evaluations were performed in an independent part of the data set. Specifically, the authors studied the effect of coding of covariables and stepwise selection on discriminative ability of the resulting model, and the effect of statistical "shrinkage" techniques on calibration. As expected, dichotomization of continuous covariables implied a loss of information. Remarkably, stepwise selection resulted in less discriminating models compared to full models including all available covariables, even when more than half of these were randomly associated with the outcome. Using qualitative information on the sign of the effect of predictors slightly improved the predictive ability. Calibration improved when shrinkage was applied on the standard maximum likelihood estimates of the regression coefficients. In conclusion, a sensible strategy in small data sets is to apply shrinkage methods in full models that include well-coded predictors that are selected based on external information.