Matching using estimated propensity scores: Relating theory to practice

Matching using estimated propensity scores: Relating theory to practice
复制标题

DOI:
10.2307/2533160
复制
发表时间:
1996-03-01
期刊:
影响因子:
1.9
通讯作者:
Thomas, N
Thomas, N
中科院分区:
数学3区
文献类型:
--
作者:
Rubin, DB;Thomas, N

文献摘要

被引文献

相似文献

配对抽样是观察性研究中评估治疗的标准技术。当有许多匹配变量时,对估计倾向分数的匹配构成了一类重要的程序。最近关于椭球分布仿射不变匹配方法的理论工作(Rubin,D.B.和Thomas,N.,1992,The Annals of Statistics 20,1079-1093)为评估这种方法的操作特性提供了一个一般框架。此外,Rubin和Thomas(1992,Bitoiska 79,797-809)使用这一框架导出了正态下匹配变量的前两个矩在样本中的分布的几个解析近似,该分布是通过对估计的线性倾向得分进行匹配而获得的。在这里,我们在这些理论近似和实际实践之间提供了一座桥梁。首先,我们完善和完善了基于正态的解析近似,从而使这些结果有可能应用于实践。其次,我们对正态和非正态椭球分布下的分析结果进行了蒙特卡罗评估,证实了解析近似的准确性,并展示了当椭球族内的正态假设被违反时,近似与模拟结果偏离的可预测方式。第三,我们将解析近似应用于具有明显非椭球分布的实际数据,并表明理论表达式尽管是在人工分布条件下导出的,但对实践具有有用的指导作用。我们的结果描绘了对估计的线性倾向分数进行匹配的广泛设置,从而为匹配研究的设计提供了有用的信息。当与特定的数据集匹配时,我们的理论近似为有利条件下的预期性能提供基准,从而识别需要特殊处理的匹配变量。在完成匹配和数据分析后,我们的结果提供了计算常见估计器的有效标准误差所需的方差。
Matched sampling is a standard technique in the evaluation of treatments in observational studies. Matching on estimated propensity scores comprises an important class of procedures when there are numerous matching variables. Recent theoretical work (Rubin, D. B. and Thomas, N., 1992, The Annals of Statistics 20, 1079-1093) on affinely invariant matching methods with ellipsoidal distributions provides a general framework for evaluationg the operating characteristics of such methods. Moreover, Rubin and Thomas (1992, Biometrika 79, 797-809) uses this framework to derive several analytic approximations under normality for the distribution of the first two moments of the matching variables in samples obtained by matching on estimated linear propensity scores. Here we provide a bridge between these theoretical approximations and actual practice. First, we complete and refine the nomal-based analytic approximations, thereby making it possible to apply these results to practice. Second, we perform Monte Carlo evaluations of the analytic results under normal and nonnormal ellipsoidal distributions, which confirm the accuracy of the analytic approximations, and demonstrate the predictable ways in which the approximations deviate from simulation results when normal assumptions are violated within the ellipsoidal family. Third, we apply the analytic approximations to real data with clearly nonellipsoidal distributions, and show that the thoretical expressions, although derived under artificial distributional conditions, produce useful guidance for practice. Our results delineate the wide range of settings in which matching on estimated Linear propensity scores performs well, thereby providing useful information for the design of matching studies. When matching with a particular data set, our theoretical approximations provide benchmarks for expected performance under favorable conditions, thereby identifying matching variables requiring special treatment. After matching is complete and data analysis is at hand, our results provide the variances required to compute valid standard errors for common estimators.