Propensity score prediction for electronic healthcare databases using super learner and high-dimensional propensity score methods

Propensity score prediction for electronic healthcare databases using super learner and high-dimensional propensity score methods
复制标题

DOI:
10.1080/02664763.2019.1582614
复制
发表时间:
2019-09-10
影响因子:
1.5
通讯作者:
van der Laan, Mark J.
van der Laan, Mark J.
中科院分区:
数学4区
文献类型:
--
作者:
Ju, Cheng;Combs, Mary;van der Laan, Mark J.

文献摘要

被引文献

相似文献

预测建模的最佳学习器根据底层数据生成分布的不同而变化。超级学习器 (SL) 是一种通用集成学习算法,它使用交叉验证在候选预测模型“库”中进行选择。虽然 SL 已在许多环境中得到广泛研究,但尚未在药物流行病学和比较有效性研究中常见的大型电子医疗保健数据库中进行彻底评估。在本研究中,我们使用三个电子医疗保健数据库应用并评估了 SL 在预测倾向评分 (PS) 的能力方面的性能,即给定基线协变量的治疗分配的条件概率。我们考虑了一个由非参数模型和参数模型组成的算法库。我们还提出了一种新的预测建模策略,将 SL 与高维倾向得分 (hdPS) 变量选择算法相结合。使用三个指标评估预测性能:负对数似然、曲线下面积 (AUC) 和时间复杂度。结果表明,就预测性能而言,最佳的单个算法因数据集而异。 SL 能够适应给定的数据集并优化相对于任何个体学习者的预测性能。将 SL 与 hdPS 相结合是最一致的预测方法,并且可能有望用于电子医疗数据库中的 PS 估计和预测建模。
The optimal learner for prediction modeling varies depending on the underlying data-generating distribution. Super Learner (SL) is a generic ensemble learning algorithm that uses cross-validation to select among a 'library' of candidate prediction models. While SL has been widely studied in a number of settings, it has not been thoroughly evaluated in large electronic healthcare databases that are common in pharmacoepidemiology and comparative effectiveness research. In this study, we applied and evaluated the performance of SL in its ability to predict the propensity score (PS), the conditional probability of treatment assignment given baseline covariates, using three electronic healthcare databases. We considered a library of algorithms that consisted of both nonparametric and parametric models. We also proposed a novel strategy for prediction modeling that combines SL with the high-dimensional propensity score (hdPS) variable selection algorithm. Predictive performance was assessed using three metrics: the negative log-likelihood, area under the curve (AUC), and time complexity. Results showed that the best individual algorithm, in terms of predictive performance, varied across datasets. The SL was able to adapt to the given dataset and optimize predictive performance relative to any individual learner. Combining the SL with the hdPS was the most consistent prediction method and may be promising for PS estimation and prediction modeling in electronic healthcare databases.