Oxford Research Encyclopedia of Economics and Finance

Oxford Research Encyclopedia of Economics and Finance
复制标题

牛津研究经济与金融百科全书

DOI:
10.1093/acrefore/9780190625979.013.256
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Kreif N
Kreif N
中科院分区:
--
文献类型:
--
作者:
Kreif N

文献摘要

相似文献

虽然机器学习(ML)方法近年来受到了很多关注,但这些方法主要用于预测。另一方面,进行政策评估的实证研究人员则专注于因果问题,试图回答反事实的问题:如果没有政策,会发生什么?由于这些反事实永远无法直接观察到(被描述为“因果推理的基本问题”),因此ML文献中的预测工具无法轻易用于因果推理。在过去的十年中,已经发生了重大创新,将监督ML工具纳入因果参数的估计器中,例如平均治疗效果(ATE)。这有望减少模型错误指定问题,并增加模型选择的透明度。一个特别成熟的文献链包括在\textit{unconfoundedness}和阳性假设(也称为交换和重叠假设)下将监督ML方法纳入二元治疗的ATE估计的方法。本文回顾了流行的监督机器学习算法,包括超级学习者。然后,介绍和说明了机器学习在治疗效果估计中的一些具体用途,即(1)在治疗组和对照组之间建立平衡,(2)估计所谓的滋扰模型(例如倾向分数,或结果的条件期望)在以因果参数为目标的半参数估计中(例如,目标最大似然估计或双ML估计),以及(3)在具有大量协变量的情况下使用机器学习进行变量选择。
While machine learning (ML) methods have received a lot of attention in recent years, these methods are primarily for prediction. Empirical researchers conducting policy evaluations are, on the other hand, pre-occupied with causal problems, trying to answer counterfactual questions: what would have happened in the absence of a policy? Because these counterfactuals can never be directly observed (described as the "fundamental problem of causal inference") prediction tools from the ML literature cannot be readily used for causal inference. In the last decade, major innovations have taken place incorporating supervised ML tools into estimators for causal parameters such as the average treatment effect (ATE). This holds the promise of attenuating model misspecification issues, and increasing of transparency in model selection. One particularly mature strand of the literature include approaches that incorporate supervised ML approaches in the estimation of the ATE of a binary treatment, under the \textit{unconfoundedness} and positivity assumptions (also known as exchangeability and overlap assumptions). This article reviews popular supervised machine learning algorithms, including the Super Learner. Then, some specific uses of machine learning for treatment effect estimation are introduced and illustrated, namely (1) to create balance among treated and control groups, (2) to estimate so-called nuisance models (e.g. the propensity score, or conditional expectations of the outcome) in semi-parametric estimators that target causal parameters (e.g. targeted maximum likelihood estimation or the double ML estimator), and (3) the use of machine learning for variable selection in situations with a high number of covariates.