A Theory of Statistical Inference for Matching Methods in Causal Research

A Theory of Statistical Inference for Matching Methods in Causal Research
复制标题

DOI:
10.1017/pan.2018.29
复制
发表时间:
2019-01-01
期刊:
影响因子:
5.4
通讯作者:
Porro, Giuseppe
Porro, Giuseppe
中科院分区:
法学1区
文献类型:
--
作者:
Iacus, Stefano M.;King, Gary;Porro, Giuseppe

文献摘要

被引文献

相似文献

生成数据的研究人员通常通过选择分层而不是简单的随机抽样设计来优化效率和稳健性。然而,为证明匹配方法的合理性而提出的所有推理理论都是基于简单的随机抽样。这更令人不安,因为尽管这些理论需要精确匹配,但大多数匹配应用程序都会诉诸某种形式的事后分层(在倾向得分、距离度量或协变量上)来寻找近似匹配,从而使这些理论旨在确保的统计特性失效。幸运的是,推理理论中使用的抽样类型是一个公理,而不是一个容易被证明错误的假设,因此我们可以用分层抽样代替简单抽样,只要我们能够像这里所做的那样,表明该理论的含义是连贯的并且保持正确。基于该理论的估计量的属性更容易理解,并且可以在没有现有理论的不吸引人的属性的情况下得到满足,例如隐藏在数据分析中而不是预先陈述的假设、渐近、不熟悉的估计量和复杂的方差计算。我们的推理理论使研究人员可以将匹配视为一种简单的预处理形式,以减少模型依赖性,之后可以应用所有熟悉的推理技术和不确定性计算。该理论还允许从一开始就使用二元、多类别和连续的治疗变量,以及针对不完美的治疗分配和不同版本的治疗的直接扩展。
Researchers who generate data often optimize efficiency and robustness by choosing stratified over simple random sampling designs. Yet, all theories of inference proposed to justify matching methods are based on simple random sampling. This is all the more troubling because, although these theories require exact matching, most matching applications resort to some form of ex post stratification (on a propensity score, distance metric, or the covariates) to find approximate matches, thus nullifying the statistical properties these theories are designed to ensure. Fortunately, the type of sampling used in a theory of inference is an axiom, rather than an assumption vulnerable to being proven wrong, and so we can replace simple with stratified sampling, so long as we can show, as we do here, that the implications of the theory are coherent and remain true. Properties of estimators based on this theory are much easier to understand and can be satisfied without the unattractive properties of existing theories, such as assumptions hidden in data analyses rather than stated up front, asymptotics, unfamiliar estimators, and complex variance calculations. Our theory of inference makes it possible for researchers to treat matching as a simple form of preprocessing to reduce model dependence, after which all the familiar inferential techniques and uncertainty calculations can be applied. This theory also allows binary, multicategory, and continuous treatment variables from the outset and straightforward extensions for imperfect treatment assignment and different versions of treatments.