Selecting Instrumental Variables in a Data Rich Environment

Selecting Instrumental Variables in a Data Rich Environment
复制标题

在数据丰富的环境中选择工具变量

DOI:
--
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
Jushan Bai
Jushan Bai
中科院分区:
--
文献类型:
--
作者:
Serena Ng;Jushan Bai

文献摘要

被引文献

相似文献

从业者通常有大量的工具可供他们使用,这些工具对于感兴趣的参数是弱外生的。然而,并不是每一种工具对内生变量都有相同的预测能力,使用太多的工具会导致偏差。我们考虑了处理这些问题的两种方法。一是从观测到的仪器中形成主成分,二是通过选择子集变量来减少仪器的数量。对于后者,我们考虑提升,一种不需要仪器先验排序的方法。我们还提出了一种预先订购仪器的方法,然后使用第一阶段回归的拟合优度和信息标准筛选仪器。我们发现主成分通常比观测数据更好,除非相关仪器的数量很少。虽然没有单一的方法占主导地位,但基于t检验的硬阈值方法通常产生的估计值具有较小的偏差和较小的均方根误差。
Practitioners often have at their disposal a large number of instruments that are weakly exogenous for the parameter of interest. However, not every instrument has the same predictive power for the endogenous variable, and using too many instruments can induce bias. We consider two ways of handling these problems. The first is to form principal components from the observed instruments, and the second is to reduce the number of instruments by subset variable selection. For the latter, we consider boosting, a method that does not require an a priori ordering of the instruments. We also suggest a way to pre-order the instruments and then screen the instruments using the goodness of fit of the first stage regression and information criteria. We find that the principal components are often better instruments than the observed data except when the number of relevant instruments is small. While no single method dominates, a hard-thresholding method based on the t test generally yields estimates with small biases and small root-mean-squared errors.