To Adjust or not to Adjust? Estimating the Average Treatment Effect in Randomized Experiments with Missing Covariates

To Adjust or not to Adjust? Estimating the Average Treatment Effect in Randomized Experiments with Missing Covariates
复制标题

DOI:
10.1080/01621459.2022.2123814
复制
发表时间:
2021-07
影响因子:
3.7
通讯作者:
Anqi Zhao;Peng Ding
Anqi Zhao;Peng Ding
中科院分区:
数学1区
文献类型:
--
作者:
Anqi Zhao;Peng Ding

文献摘要

相似文献

摘要 随机实验允许根据平均结果的差异对平均治疗效果进行一致的估计,而无需强有力的建模假设。适当使用预处理协变量可以进一步提高估计效率。然而,协变量的缺失在实践中很常见,并提出了一个重要的问题:我们是否应该针对缺失的协变量进行调整?如果是,如何调整?未经调整的均值差异始终是无偏的。完整协变量分析针对所有完全观察到的协变量进行调整,并且如果至少一个完全观察到的协变量可以预测结果,则它比均值差异渐近更有效。那么调整协变量的额外增益是多少呢?为了调和文献中相互矛盾的建议,我们分析和比较了基于设计的框架下随机实验中处理缺失协变量的五种策略,并推荐缺失指示符方法,由于其多重优点,作为文献中已知但不那么流行的策略。首先,它消除了回归调整估计量对缺失协变量的估算值的依赖性。其次,它不需要对缺失机制进行建模,即使缺失机制与缺失的协变量和不可观察的潜在结果相关,也能产生一致的估计量。第三,它确保了完整协变量分析和仅基于估算协变量的分析的大样本效率。最后,通过最小二乘法很容易实现。我们还基于渐近和有限样本的考虑提出对其进行修改。重要的是,我们的理论将随机化视为推理的基础,并且不对数据生成过程或缺失机制强加任何建模假设。本文的补充材料可在线获取。
Abstract Randomized experiments allow for consistent estimation of the average treatment effect based on the difference in mean outcomes without strong modeling assumptions. Appropriate use of pretreatment covariates can further improve the estimation efficiency. Missingness in covariates is nevertheless common in practice, and raises an important question: should we adjust for covariates subject to missingness, and if so, how? The unadjusted difference in means is always unbiased. The complete-covariate analysis adjusts for all completely observed covariates, and is asymptotically more efficient than the difference in means if at least one completely observed covariate is predictive of the outcome. Then what is the additional gain of adjusting for covariates subject to missingness? To reconcile the conflicting recommendations in the literature, we analyze and compare five strategies for handling missing covariates in randomized experiments under the design-based framework, and recommend the missingness-indicator method, as a known but not so popular strategy in the literature, due to its multiple advantages. First, it removes the dependence of the regression-adjusted estimators on the imputed values for the missing covariates. Second, it does not require modeling the missingness mechanism, and yields consistent estimators even when the missingness mechanism is related to the missing covariates and unobservable potential outcomes. Third, it ensures large-sample efficiency over the complete-covariate analysis and the analysis based on only the imputed covariates. Lastly, it is easy to implement via least squares. We also propose modifications to it based on asymptotic and finite sample considerations. Importantly, our theory views randomization as the basis for inference, and does not impose any modeling assumptions on the data-generating process or missingness mechanism. Supplementary materials for this article are available online.