Too many covariates and too few cases? - a comparative study

Too many covariates and too few cases? - a comparative study
复制标题

DOI:
10.1002/sim.7021
复制
发表时间:
2016-10-01
影响因子:
2
通讯作者:
Harrell, Frank E., Jr.
Harrell, Frank E., Jr.
中科院分区:
医学3区
文献类型:
--
作者:
Chen, Qingxia;Nian, Hui;Harrell, Frank E., Jr.

文献摘要

被引文献

相似文献

先前的研究表明,在多变量逻辑回归模型中,每个参数需要10-15个病例或对照,以较少者为准,才能可靠地估计回归系数。即使在设计良好的研究中,当潜在混杂因素的数量很大,结果很少,和/或相互作用很有趣时,这种情况也可能难以满足。当暴露为二值时,已经实施了各种倾向评分方法。最近关于收缩方法(如lasso)的工作是由迫切需要开发p >> n情况的方法所驱动的,其中p是参数的数量,n是样本量。然而,当p近似于n时,这些方法的使用频率较低,在这种情况下,没有关于在规则逻辑回归模型、倾向评分方法和收缩方法之间进行选择的指导。为了填补这一空白,我们进行了广泛的模拟,模拟我们的临床数据,估计疫苗在2011-2012年流感季节预防流感住院的有效性。在这些类型的研究中可以考虑岭回归和惩罚逻辑回归模型,它们惩罚除了暴露系数之外的所有因素。版权所有:John Wiley & Sons, Ltd。
Prior research indicates that 10-15 cases or controls, whichever fewer, are required per parameter to reliably estimate regression coefficients in multivariable logistic regression models. This condition may be difficult tomeet even in a well-designed study when the number of potential confounders is large, the outcome is rare, and/or interactions are of interest. Various propensity score approaches have been implemented when the exposure is binary. Recent work on shrinkage approaches like lasso were motivated by the critical need to develop methods for the p >> n situation, where p is the number of parameters and n is the sample size. Those methods, however, have been less frequently used when p approximate to n, and in this situation, there is no guidance on choosing among regular logistic regression models, propensity score methods, and shrinkage approaches. To fill this gap, we conducted extensive simulations mimicking our motivating clinical data, estimating vaccine effectiveness for preventing influenza hospitalizations in the 2011-2012 influenza season. Ridge regression and penalized logistic regression models that penalize all but the coefficient of the exposure may be considered in these types of studies. Copyright (C) 2016 John Wiley & Sons, Ltd.