Overlap in observational studies with high-dimensional covariates

Overlap in observational studies with high-dimensional covariates
复制标题

DOI:
10.1016/j.jeconom.2019.10.014
复制
发表时间:
2021-02-11
影响因子:
6.3
通讯作者:
Sekhon, Jasjeet
Sekhon, Jasjeet
中科院分区:
经济学2区
文献类型:
--
作者:
D'Amour, Alexander;Ding, Peng;Sekhon, Jasjeet

文献摘要

被引文献

相似文献

估计外生性下的因果效应取决于两个关键假设:无混性和重叠性。研究人员经常争辩说,当分析中包括更多的协变量时,不混淆更有可能。较少讨论的事实是,在这种情况下,协变量重叠更难满足。在这篇文章中,我们探索了高维协变量观察研究中重叠的含义,并将维度灾难论点形式化,表明这些假设比研究人员可能意识到的更强。我们的关键创新是探索严格重叠如何限制治疗人群和对照人群中协变量分布之间的全局差异。利用信息论的结果,我们得到了严格重叠下协变量均值的平均不平衡的显式界,并且表明,随着维度的增大,这些界变得更具限制性。我们讨论了这些含义如何与观测因果推理中常见的假设和过程相互作用,包括稀疏性和剪裁。(C)2020作者。爱思唯尔出版公司(Elsevier B.V.)
Estimating causal effects under exogeneity hinges on two key assumptions: unconfoundedness and overlap. Researchers often argue that unconfoundedness is more plausible when more covariates are included in the analysis. Less discussed is the fact that covariate overlap is more difficult to satisfy in this setting. In this paper, we explore the implications of overlap in observational studies with high-dimensional covariates and formalize curse-of-dimensionality argument, suggesting that these assumptions are stronger than investigators likely realize. Our key innovation is to explore how strict overlap restricts global discrepancies between the covariate distributions in the treated and control populations. Exploiting results from information theory, we derive explicit bounds on the average imbalance in covariate means under strict overlap and show that these bounds become more restrictive as the dimension grows large. We discuss how these implications interact with assumptions and procedures commonly deployed in observational causal inference, including sparsity and trimming. (C) 2020 The Authors. Published by Elsevier B.V.