Estimating causal effects from large data sets using propensity scores

Estimating causal effects from large data sets using propensity scores
复制标题

DOI:
10.7326/0003-4819-127-8_part_2-199710151-00064
复制
发表时间:
1997-10-15
影响因子:
39.2
通讯作者:
Rubin, DB
Rubin, DB
中科院分区:
医学1区
文献类型:
--
作者:
Rubin, DB

文献摘要

被引文献

相似文献

对大型数据库进行诸多分析的目的是就行动、治疗或干预措施的效果得出因果推论。例如,医生治疗某一特定患者时可用的各种选择的效果、不同医疗服务提供者的相对疗效,以及实施一项新的国家医疗保健政策的后果。利用大型数据库实现这些目的的一个复杂之处在于,其数据几乎总是观察性的,而非实验性的。也就是说,大多数大型数据集中的数据并非基于精心开展的随机临床试验的结果,而是代表在正常实践中对系统进行观察所收集的数据,没有按照随机分配规则实施任何干预措施。然而,获取此类数据相对成本较低,而且往往比随机实验环境更能代表医疗实践的全貌。因此,尝试从这类大型数据集中估计治疗效果是明智的,即使只是为了帮助设计一项新的随机实验,或者阐明现有随机实验结果的普遍性。然而,使用现有的统计软件(如线性回归或逻辑回归)的标准分析方法对于这些目标可能具有误导性,因为它们没有就其合理性给出警示。倾向评分方法是实现这些目标更可靠的工具,因为使这些方法得出的答案合理所需的假设对研究者来说更易于评估且更透明。
The aim of many analyses of large databases is to draw causal inferences about the effects of actions, treatments, or interventions. Examples include the effects of various options available to a physician for treating a particular patient, the relative efficacies of various health care providers, and the consequences of implementing a new national health care policy. A complication of using large databases to achieve such aims is that their data are almost always observational rather than experimental. That is, the data in most large data sets are not based on the results of carefully conducted randomized clinical trials, but rather represent data collected through the observation of systems as they operate in normal practice without any interventions implemented by randomized assignment rules. Such data are relatively inexpensive to obtain, however, and often do represent the spectrum of medical practice better than the settings of randomized experiments. Consequently, it is sensible to try to estimate the effects of treatments from such large data sets, even if only to help design a new randomized experiment or shed light on the generalizability of results from existing randomized experiments. However, standard methods of analysis using available statistical software (such as linear or logistic regression) can be deceptive for these objectives because they provide no warnings about their propriety. Propensity score methods are more reliable tools for addressing such objectives because the assumptions needed to make their answers appropriate are more assessable and transparent to the investigator.