Automated versus Do-It-Yourself Methods for Causal Inference: Lessons Learned from a Data Analysis Competition

Automated versus Do-It-Yourself Methods for Causal Inference: Lessons Learned from a Data Analysis Competition
复制标题

DOI:
10.1214/18-sts667
复制
发表时间:
2019-02-01
影响因子:
5.7
通讯作者:
Cervone, Dan
Cervone, Dan
中科院分区:
数学2区
文献类型:
--
作者:
Dorie, Vincent;Hill, Jennifer;Cervone, Dan

文献摘要

被引文献

相似文献

统计学家在创造减少我们对参数假设的依赖的方法方面取得了很大进展。然而,这种爆炸性的研究导致了广泛的推理策略,既创造了更可靠的推理机会,也使应用研究人员必须做出和捍卫的选择复杂化。与此相关的是,倡导新方法的研究人员通常将他们的方法与最多2或3种其他因果推理策略进行比较,并使用模拟进行测试,这些模拟可能会或可能不会被设计为同样梳理出所有竞争方法中的缺陷。因果推理数据分析挑战,“你的SATT在哪里?“,作为2016年大西洋因果推理会议的一部分,寻求在这两个问题上取得进展。创建数据测试场的研究人员与提交其有效性将被评估的方法的研究人员不同。来自30个竞争对手的两个版本的竞争(黑盒算法和自己动手的分析)的结果与事后分析,揭示信息的因果推理策略和设置的特点,影响性能的沿着。最一致的结论是,灵活地模拟响应面的方法比未能这样做的方法整体表现更好。最后,提出了新的方法,联合收割机功能的几个顶级性能提交的方法。
Statisticians have made great progress in creating methods that reduce our reliance on parametric assumptions. However, this explosion in research has resulted in a breadth of inferential strategies that both create opportunities for more reliable inference as well as complicate the choices that an applied researcher has to make and defend. Relatedly, researchers advocating for new methods typically compare their method to at best 2 or 3 other causal inference strategies and test using simulations that may or may not be designed to equally tease out flaws in all the competing methods. The causal inference data analysis challenge, "Is Your SATT Where It's At?", launched as part of the 2016 Atlantic Causal Inference Conference, sought to make progress with respect to both of these issues. The researchers creating the data testing grounds were distinct from the researchers submitting methods whose efficacy would be evaluated. Results from 30 competitors across the two versions of the competition (black-box algorithms and do-it-yourself analyses) are presented along with post-hoc analyses that reveal information about the characteristics of causal inference strategies and settings that affect performance. The most consistent conclusion was that methods that flexibly model the response surface perform better overall than methods that fail to do so. Finally new methods are proposed that combine features of several of the top-performing submitted methods.