Dummy Variables: Mechanics v. Interpretation

Dummy Variables: Mechanics v. Interpretation
复制标题

虚拟变量:力学与解释

DOI:
--
复制
发表时间:
1984
期刊:
影响因子:
--
通讯作者:
D. Suits
D. Suits
中科院分区:
--
文献类型:
--
作者:
D. Suits

文献摘要

被引文献

相似文献

包含虚拟变量的回归很容易通过熟悉的权宜之计“放弃”其中一个类别来估计,但结果往往难以解释。然而,由于虚拟变量的系数仅确定到一个附加常数,因此可以通过将适当选择的常数添加到每个系数来将方程转换为更容易解释的形式。对于大多数回归,应选择常数以迫使变换系数的平均值等于0。对于对数回归,常数的选择应使系数的反对数之和等于1。用对数需求曲线拟合月度数据,得到的反对数就成为月度季节指数。在回归方程中使用虚拟变量来捕捉分类变量的影响的技术程序通常是熟悉的(参见Goldberger(1964),Kmenta(1971),约翰斯顿(1960),或者,回到事情的开始,Suits(1957))。在许多情况下,特别是在只涉及两类观察的情况下,以通常方式提出的结果不涉及解释的特殊问题。例如,使用虚拟变量来区分战前和战后的行为,或衡量罢工期间关系的变化,任何读者都很容易理解。但是,如果使用一组虚拟变量来衡量不同阶层(地区、教育群体、年龄段等)之间的行为差异,那么在拟合回归的纯机械问题和以最有效的方式呈现结果的完全不同的问题之间,往往存在着重要的区别。本文的目的是提请注意这种区别,并通过简单的例子来说明。
Regressions containing dummy variables are easily estimated by the familiar expedient of "dropping out" one of the categories but the result is often awkward to interpret. Since coefficients of dummy variables are determined only up to an additive constant, however, the equation can be transformed into a more easily interpretable form by adding on an appropriately chosen constant to each coefficient. For most regressions the constants should be chosen to force the mean of the transformed coefficients to equal 0. For logarithmic regressions the constants should be chosen to force the sum of the antilogs of the coefficients to equal 1. With logarithmic demand curves fitted to monthly data the resulting antilogs become monthly seasonal indexes. The technical procedure by which dummy variables are used to capture the influence of categorical variables in regression equations is generally familiar (see Goldberger (1964), Kmenta (1971), Johnston (1960), or, to go back near the beginning of things, Suits (1957)). In many cases, particularly where only two classes of observation are involved, results presented in the usual way involve no special problems of interpretation. For example, use of a dummy variable to distinguish pre-war from post-war behavior, or to measure the shift in a relationship during the period of a strike is readily understood by any reader. But where a set of several dummy variables is employed to measure the variation in behavior among a number of classes-regions, education groups, age brackets, and the like-there is often an important difference between the purely mechanical problem of fitting the regression and the quite different problem of presenting the results in the most effective fashion. The purpose of this paper is to call attention to this distinction, and to illustrate by simple examples.