Optimizing Variance-Bias Trade-off in the TWANG Package for Estimation of Propensity Scores.

Optimizing Variance-Bias Trade-off in the TWANG Package for Estimation of Propensity Scores.
复制标题

DOI:
10.1007/s10742-016-0168-2
复制
发表时间:
2017-12
影响因子:
1.5
通讯作者:
Griffin BA
Griffin BA
中科院分区:
其他
文献类型:
--
作者:
Parast L;McCaffrey DF;Burgette LF;de la Guardia FH;Golinelli D;Miles JNV;Griffin BA

文献摘要

被引文献

相似文献

虽然倾向得分加权已被证明可以减少存在选择偏差时治疗效果估计中的偏差,但也已表明,如果估计的倾向得分权重变化很大,则这种加权可能表现不佳。人们已经提出了各种方法来减少权重的可变性和性能不佳的风险,特别是那些基于机器学习方法的方法。在这项研究中,我们仔细研究了微调一种机器学习技术(广义增强模型 [GBM])的方法,以选择倾向得分,以寻求优化大多数倾向得分分析中固有的方差-偏差权衡。具体来说,我们提出并评估了三种方法,用于在 R 的 twang 包中选择 GBM 的最佳树数。通常,R 中的 twang 包会迭代地选择最佳树数,以最大化所考虑的处理组之间的平衡。由于所选的树木数量可能会导致倾向得分权重高度可变,因此我们研究了调整倾向得分权重估计中使用的树木数量的替代方法,以便我们牺牲预处理协变量的一些平衡,以换取较少的可变权重。我们使用模拟研究来说明这些方法并描述每种方法的潜在优点和缺点。我们将这些方法应用于两个案例研究:一个使用来自加利福尼亚州一项大型人口调查的数据来研究养狗对主人总体健康的影响,第二个研究在高风险青少年样本中禁欲与长期经济结果之间的关系。
While propensity score weighting has been shown to reduce bias in treatment effect estimation when selection bias is present, it has also been shown that such weighting can perform poorly if the estimated propensity score weights are highly variable. Various approaches have been proposed which can reduce the variability of the weights and the risk of poor performance, particularly those based on machine learning methods. In this study, we closely examine approaches to fine-tune one machine learning technique (generalized boosted models [GBM]) to select propensity scores that seek to optimize the variance-bias trade-off that is inherent in most propensity score analyses. Specifically, we propose and evaluate three approaches for selecting the optimal number of trees for the GBM in the twang package in R. Normally, the twang package in R iteratively selects the optimal number of trees as that which maximizes balance between the treatment groups being considered. Because the selected number of trees may lead to highly variable propensity score weights, we examine alternative ways to tune the number of trees used in the estimation of propensity score weights such that we sacrifice some balance on the pre-treatment covariates in exchange for less variable weights. We use simulation studies to illustrate these methods and to describe the potential advantages and disadvantages of each method. We apply these methods to two case studies: one examining the effect of dog ownership on the owner’s general health using data from a large, population-based survey in California, and a second investigating the relationship between abstinence and a long-term economic outcome among a sample of high-risk youth.