Estimating predicted probabilities from logistic regression: different methods correspond to different target populations

Estimating predicted probabilities from logistic regression: different methods correspond to different target populations
复制标题

DOI:
10.1093/ije/dyu029
复制
发表时间:
2014-06-01
影响因子:
7.7
通讯作者:
MacLehose, Richard F.
MacLehose, Richard F.
中科院分区:
医学1区
文献类型:
--
作者:
Muller, Clemma J.;MacLehose, Richard F.

文献摘要

被引文献

相似文献

背景:我们回顾了三种常用的估计混杂因素调整后logistic回归预测概率的方法:边际标准化预测概率求和为反映目标人群中混杂因素分布的加权平均值);在模态下的预测(条件预测概率通过将每个混杂因素设置为其模态值来计算);均值预测通过将每个混杂因素设置为其均值来预测概率)。每种方法对应于不同的目标人群,这一点在实践中没有得到充分的重视。具体来说,平均预测常常被错误地解释为对整个研究群体的平均概率的估计,并且在存在二分类混杂因素的情况下产生无意义的估计。流行统计软件包中的默认命令经常导致无意中误用均值预测。方法:通过一个应用实例,我们展示了这些方法在预测概率上的差异,讨论了解释的含义,并提供了SAS和Stata的语法。结果:边际标准化允许对抽取数据的总体人口进行推断。在模态或均值上的预测只允许对观测的相关层进行推断。对于二分类混杂因素,平均预测对应于不包含任何实际观察的地层。结论:在对总体人群进行推断时,边际标准化是合适的方法。其他方法应谨慎使用,均值预测不应与二元混杂因素一起使用。Stata,而不是SAS,采用了简单的边际标准化方法。
Background: We review three common methods to estimate predicted probabilities following confounder-adjusted logistic regression: marginal standardization predicted probabilities summed to a weighted average reflecting the confounder distribution in the target population); prediction at the modes conditional predicted probabilities calculated by setting each confounder to its modal value); and prediction at the means predicted probabilities calculated by setting each confounder to its mean value). That each method corresponds to a different target population is underappreciated in practice. Specifically, prediction at the means is often incorrectly interpreted as estimating average probabilities for the overall study population, and furthermore yields nonsensical estimates in the presence of dichotomous confounders. Default commands in popular statistical software packages often lead to inadvertent misapplication of prediction at the means.Methods: Using an applied example, we demonstrate discrepancies in predicted probabilities across these methods, discuss implications for interpretation and provide syntax for SAS and Stata.Results: Marginal standardization allows inference to the total population from which data are drawn. Prediction at the modes or means allows inference only to the relevant stratum of observations. With dichotomous confounders, prediction at the means corresponds to a stratum that does not include any real-life observations.Conclusions: Marginal standardization is the appropriate method when making inference to the overall population. Other methods should be used with caution, and prediction at the means should not be used with binary confounders. Stata, but not SAS, incorporates simple methods for marginal standardization.