Scaling regression inputs by dividing by two standard deviations

Scaling regression inputs by dividing by two standard deviations
复制标题

DOI:
10.1002/sim.3107
复制
发表时间:
2008-07-10
影响因子:
2
通讯作者:
Gelman, Andrew
Gelman, Andrew
中科院分区:
医学3区
文献类型:
--
作者:
Gelman, Andrew

文献摘要

被引文献

相似文献

回归系数的解释对输入的规模很敏感。一种常用于将输入变量置于通用尺度上的方法是将每个数值变量除以其标准差。在这里,我们建议将每个数字变量除以其标准差的两倍,以便通用比较的输入等于平均值+/- 1标准差。然后,所得系数直接与未变换的二进制预测因子相当。我们已经将该过程实现为R中的函数。我们用两个简单的分析来说明这种方法,这两个分析是典型的应用模型:全国选举研究数据的线性回归和纽约市公寓啮齿动物患病率数据的多层次逻辑回归。我们建议将重新封装作为默认选项,这是对通常方法的改进,即以数据文件中的任何编码方式包含变量,以便可以将系数的大小直接比较作为常规统计实践。版权所有(C)2007约翰威利父子有限公司
Interpretation of regression coefficients is sensitive to the scale of the inputs. One method often used to place input variables on a common scale is to divide each numeric variable by its standard deviation. Here we propose dividing each numeric variable by two times its standard deviation, so that the generic comparison is with inputs equal to the mean +/- 1 standard deviation. The resulting coefficients are then directly comparable for untransformed binary predictors. We have implemented the procedure as a function in R. We illustrate the method with two simple analyses that are typical of applied modeling: a linear regression of data from the National Election Study and a multilevel logistic regression of data on the prevalence of rodents in New York City apartments. We recommend our resealing as a default option-an improvement upon the usual approach of including variables in whatever way they are coded in the data file-so that the magnitudes of coefficients can be directly compared as a matter of routine statistical practice. Copyright (C) 2007 John Wiley & Sons, Ltd.