Hypothesis Tests That Are Robust to Choice of Matching Method

Hypothesis Tests That Are Robust to Choice of Matching Method
复制标题

对匹配方法的选择具有鲁棒性的假设检验

DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
C. Rudin
C. Rudin
中科院分区:
--
文献类型:
--
作者:
Marco Morucci;Md. Noor;C. Rudin

文献摘要

被引文献

相似文献

大量的因果推理研究在治疗病例与类似的对照病例相匹配后测试关于治疗效果的假设。匹配数据的质量通常根据一些度量标准进行评估,例如平衡;然而,相同数据上的不同匹配可以实现相同级别的匹配质量。至关重要的是,达到相同质量水平的匹配可能会导致对匹配数据进行假设检验的不同结果。实验者通常会特别选择不考虑由匹配构建方式产生的不确定性;这使得计算更容易,测试更清晰,但它不会考虑分配构建方式中可能存在的偏差。我们真正希望能够报告的是,无论我们选择哪种分配,只要匹配足够好,那么假设检验结果仍然成立。在本文中,我们提供了基于离散优化的方法,以创建明确说明这种变化的强大的测试。对于二进制数据,我们给出了两个快速算法来计算我们的测试和公式的零分布下,我们的测试统计量的不同概念的匹配。对于连续数据,我们制定了一个强大的测试统计量,并提供了一个线性化,允许更快的计算。我们将我们的方法应用于现实世界的数据集,并表明它们可以在实际应用环境中产生有用的结果。
A vast number of causal inference studies test hypotheses on treatment effects after treatment cases are matched with similar control cases. The quality of matched data is usually evaluated according to some metric, such as balance; however the same level of match quality can be achieved by different matches on the same data. Crucially, matches that achieve the same level of quality might lead to different results for hypothesis tests conducted on the matched data. Experimenters often specifically choose not to consider the uncertainty stemming from how the matches were constructed; this allows for easier computation and clearer testing, but it does not consider possible biases in the way the assignments were constructed. What we would really like to be able to report is that no matter which assignment we choose, as long as the match is sufficiently good, then the hypothesis test result still holds. In this paper, we provide methodology based on discrete optimization to create robust tests that explicitly account for this variation. For binary data, we give both fast algorithms to compute our tests and formulas for the null distributions of our test statistics under different conceptions of matching. For continuous data, we formulate a robust test statistic, and offer a linearization that permits faster computation. We apply our methods to real-world datasets and show that they can produce useful results in practical applied settings.