Estimating and using propensity scores with partially missing data

Estimating and using propensity scores with partially missing data
复制标题

DOI:
10.2307/2669455
复制
发表时间:
2000-09-01
影响因子:
3.7
通讯作者:
Rubin, DB
Rubin, DB
中科院分区:
数学1区
文献类型:
--
作者:
D'Agostino, RB;Rubin, DB

文献摘要

被引文献

相似文献

观察性研究中的研究者无法控制治疗分配。因此,治疗组和对照组之间观察到的协变量可能存在较大差异,这可能导致治疗效应的严重偏倚估计。倾向评分方法是一种越来越流行的方法,用于平衡两组中协变量的分布,以减少这种偏倚;例如,使用匹配或子分类,有时结合基于模型的调整。为了估计倾向评分,即在给定观测协变量向量的情况下接受治疗的条件概率,我们必须在给定这些观测协变量的情况下对治疗指标的分布进行建模。在协变量完全观测的情况下,已经做了很多工作。我们解决的问题,计算倾向分数时,协变量可以有缺失值。在这种情况下,这通常出现在实践中,缺失协变量的模式可能是非常重要的,然后倾向分数应该条件下观察到的协变量值和观察到的缺失数据指标。使用所得的广义倾向评分调整治疗组和对照组之间观察到的背景差异,预期会导致治疗组和对照组中观察到的协变量的平衡分布,以及缺失数据模式的平衡分布。说明使用广义倾向分数的方法,以创建匹配的样本中的过期妊娠的影响的研究。
Investigators in observational studies have no control over treatment assignment. As a result, large differences can exist between the treatment and control groups on observed covariates, which can lead to badly biased estimates of treatment effects. Propensity score methods are an increasingly popular method for balancing the distribution of the covariates in the two groups to reduce this bias; for example, using matching or subclassification, sometimes in combination with model-based adjustment. To estimate propensity scores, which are the conditional probabilities of being treated given a vector of observed covariates, we must model the distribution of the treatment indicator given these observed covariates. Much work has been done in the case where covariates are fully observed. We address the problem of calculating propensity scores when covariates can have missing values. In such cases, which commonly arise in practice, the pattern of missing covariates can be prognostically important, and then propensity scores should condition both on observed values of covariates and on the observed missing-data indicators. Using the resulting generalized propensity scores to adjust for the observed background differences between treatment and control groups leads, in expectation, to balanced distributions of observed covariates in the treatment and control groups, as well as balanced distributions of patterns of missing data. The methods are illustrated using the generalized propensity scores to create matched samples in a study of the effects of postterm pregnancy.