A comparative investigation of methods for logistic regression with separated or nearly separated data

A comparative investigation of methods for logistic regression with separated or nearly separated data
复制标题

DOI:
10.1002/sim.2687
复制
发表时间:
2006-12-30
影响因子:
2
通讯作者:
Heinze, Georg
Heinze, Georg
中科院分区:
医学3区
文献类型:
--
作者:
Heinze, Georg

文献摘要

被引文献

相似文献

在小数据集或稀疏数据集的逻辑回归分析中,经典的最大似然方法得到的结果通常不可信。在这样的分析中,甚至可能发生可能性满足收敛标准,而至少一个参数估计值发散到+/-无穷大。这种情况被称为“分离”,通常发生在由二分协变量定义的两组之一中没有观察到事件时。更一般地说,分离是由连续或二分协变量的线性组合引起的,它完美地将事件与非事件分开。分离意味着比值比的最大似然估计为无穷大或零,这通常被认为是不现实的。我提供了一些分离和接近分离的临床数据集的例子,并讨论了一些选择来分析这些数据,包括精确的逻辑回归分析和惩罚似然方法。这两种方法在分离的情况下都提供有限点估计。参数的轮廓惩罚似然置信区间在覆盖概率方面表现出优异的性能,并提供比精确置信区间更高的功效。惩罚似然法的一般优点进行了讨论。版权所有(c)2006约翰威利父子有限公司。
In logistic regression analysis of small or sparse data sets, results obtained by classical maximum likelihood methods cannot be generally trusted. In such analyses it may even happen that the likelihood meets the convergence criteria while at least one parameter estimate diverges to +/-infinity. This situation has been termed,'separation', and it typically occurs whenever no events are observed in one of the two groups defined by a dichotomous covariate. More generally, separation is caused by a linear combination of continuous or dichotomous covariates that perfectly separates events from non-events. Separation implies infinite or zero maximum likelihood estimates of odds ratios, which are usually considered unrealistic. I provide some examples of separation and near-separation in clinical data sets and discuss some options to analyse such data, including exact logistic regression analysis and a penalized likelihood approach. Both methods supply finite point estimates in case of separation. Profile penalized likelihood confidence intervals for parameters show excellent behaviour in terms of coverage probability and provide higher power than exact confidence intervals. General advantages of the penalized likelihood approach are discussed. Copyright (c) 2006 John Wiley & Sons, Ltd.