Dirichlet process mixture models to impute missing predictor data in counterfactual prediction models: an application to predict optimal type 2 diabetes therapy.

Dirichlet process mixture models to impute missing predictor data in counterfactual prediction models: an application to predict optimal type 2 diabetes therapy.
复制标题

DOI:
10.1186/s12911-023-02400-3
复制
发表时间:
2024-01-08
影响因子:
3.5
通讯作者:
--
中科院分区:
医学3区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

缺失数据的处理对于推理和回归建模来说是一个挑战。一个特殊的挑战是处理丢失的预测信息,特别是在尝试构建用于临床实践的模型并进行预测时。我们利用灵活的贝叶斯方法来处理回归模型中缺失的预测变量信息。这为从业者提供了缺失的预测变量信息(以观察到的预测变量为条件)和感兴趣的结果的完整后验预测分布。我们将此方法应用于先前提出的 2 型糖尿病二线治疗的反事实治疗选择模型。我们的方法结合了回归模型和狄利克雷过程混合模型(DPMM),其中前者定义了治疗选择模型,后者提供了一种灵活的方法来对预测变量的联合分布进行建模。我们表明,DPMM 可以对预测变量之间的复杂关系进行建模,并且可以提供将模型拟合到不完整数据的强大方法(在完全随机缺失和随机缺失假设下)。该框架确保参数的后验分布和条件平均治疗效果估计自动反映由于分层模型结构而与缺失数据相关的额外不确定性。我们还证明,在存在多个缺失预测变量的情况下,DPMM 模型可用于探索哪些变量(如果收集)可以提供有关可能结果的最多附加信息。在开发临床预测模型时,DPMM 提供了一种灵活的方法来建模复杂的协变量结构并处理缺失的预测信息。基于 DPMM 的反事实预测模型还可以提供额外的信息来支持临床决策,包括允许对预测数据不完整的个体进行具有适当不确定性的预测。在线版本包含可在 10.1186/s12911-023-02400-3 获取的补充材料。
The handling of missing data is a challenge for inference and regression modelling. A particular challenge is dealing with missing predictor information, particularly when trying to build and make predictions from models for use in clinical practice. We utilise a flexible Bayesian approach for handling missing predictor information in regression models. This provides practitioners with full posterior predictive distributions for both the missing predictor information (conditional on the observed predictors) and the outcome-of-interest. We apply this approach to a previously proposed counterfactual treatment selection model for type 2 diabetes second-line therapies. Our approach combines a regression model and a Dirichlet process mixture model (DPMM), where the former defines the treatment selection model, and the latter provides a flexible way to model the joint distribution of the predictors. We show that DPMMs can model complex relationships between predictor variables and can provide powerful means of fitting models to incomplete data (under missing-completely-at-random and missing-at-random assumptions). This framework ensures that the posterior distribution for the parameters and the conditional average treatment effect estimates automatically reflect the additional uncertainties associated with missing data due to the hierarchical model structure. We also demonstrate that in the presence of multiple missing predictors, the DPMM model can be used to explore which variable(s), if collected, could provide the most additional information about the likely outcome. When developing clinical prediction models, DPMMs offer a flexible way to model complex covariate structures and handle missing predictor information. DPMM-based counterfactual prediction models can also provide additional information to support clinical decision-making, including allowing predictions with appropriate uncertainty to be made for individuals with incomplete predictor data. The online version contains supplementary material available at 10.1186/s12911-023-02400-3.
DOI: 10.1080/01621459.2016.1231612
发表时间: 2017-01-01
影响因子: 3.7
作者:
Manrique-Vallier, Daniel;Reiter, Jerome P.
通讯作者: Reiter, Jerome P.
DOI: 10.1080/10618600.2016.1172487
发表时间: 2017-01-01
影响因子: 2.4
作者:
de Valpine, Perry;Turek, Daniel;Bodik, Rastislav
通讯作者: Bodik, Rastislav
DOI: 10.1016/j.jmp.2019.04.004
发表时间: 2019-08-01
影响因子: 1.8
作者:
Li, Yuelin;Schofield, Elizabeth;Gonen, Mithat
通讯作者: Gonen, Mithat
DOI: 10.1007/s11222-006-5196-2
发表时间: 2006-03-01
影响因子: 2.2
作者:
McAuliffe, JD;Blei, DM;Jordan, MI
通讯作者: Jordan, MI
DOI: 10.1214/aos/1176342360
发表时间: 1973-01-01
影响因子: 4.5
作者:
FERGUSON, TS
通讯作者: FERGUSON, TS