Generic Machine Learning Inference on Heterogenous Treatment Effects in Randomized Experiments

Generic Machine Learning Inference on Heterogenous Treatment Effects in Randomized Experiments
复制标题

DOI:
10.1920/wp.cem.2017.6117
复制
发表时间:
2017-12
期刊:
Randomized Social Experiments eJournal
影响因子:
--
通讯作者:
V. Chernozhukov;Mert Demirer;E. Duflo;Iván Fernández-Val
V. Chernozhukov;Mert Demirer;E. Duflo;Iván Fernández-Val
中科院分区:
其他
文献类型:
--
作者:
V. Chernozhukov;Mert Demirer;E. Duflo;Iván Fernández-Val

文献摘要

被引文献

相似文献

我们提出了在随机实验中估计和推断异质性效应关键特征的策略。这些关键特征包括使用机器学习代理的效果的最佳线性预测器,按影响组排序的平均效果,以及受影响最大和最小单位的平均特征。该方法在高维环境中是有效的,其中的效果由机器学习方法代替。我们将这些代理后处理成关键特征的估计。我们的方法是通用的,它可以与惩罚方法、深度和浅神经网络、规范和新的随机森林、增强树和集成方法结合使用。我们的方法是不可知论的,不会做出不切实际或难以检验的假设;我们不需要ML方法一致性的条件。估计和推理依赖于重复的数据分割来避免过拟合并达到有效性。对于推理,我们取许多不同数据分割产生的p值和置信区间的中位数,然后调整它们的标称水平以保证一致的有效性。这种变分推理方法是一致有效的,并且量化了参数估计和数据分裂带来的不确定性。推理方法可能在许多机器学习应用中具有实质性的独立兴趣。对小额信贷对经济发展影响的实证应用说明了该方法在随机实验中的使用。性别歧视对工资影响的另一个应用说明了该方法在观察性研究中的潜在用途,在观察性研究中,机器学习方法可用于灵活地调节非常高维的控制。
We propose strategies to estimate and make inference on key features of heterogeneous effects in randomized experiments. These key features include best linear predictors of the effects using machine learning proxies, average effects sorted by impact groups, and average characteristics of most and least impacted units. The approach is valid in high dimensional settings, where the effects are proxied by machine learning methods. We post-process these proxies into the estimates of the key features. Our approach is generic, it can be used in conjunction with penalized methods, deep and shallow neural networks, canonical and new random forests, boosted trees, and ensemble methods. Our approach is agnostic and does not make unrealistic or hard-to-check assumptions; we don’t require conditions for consistency of the ML methods. Estimation and inference relies on repeated data splitting to avoid overfitting and achieve validity. For inference, we take medians of p-values and medians of confidence intervals, resulting from many different data splits, and then adjust their nominal level to guarantee uniform validity. This variational inference method is shown to be uniformly valid and quantifies the uncertainty coming from both parameter estimation and data splitting. The inference method could be of substantial independent interest in many machine learning applications. An empirical application to the impact of micro-credit on economic development illustrates the use of the approach in randomized experiments. An additional application to the impact of the gender discrimination on wages illustrates the potential use of the approach in observational studies, where machine learning methods can be used to condition flexibly on very high-dimensional controls.