Dirichlet process mixture models to estimate outcomes for individuals with missing predictor data: application to predict optimal type 2 diabetes therapy in electronic health record data

Dirichlet process mixture models to estimate outcomes for individuals with missing predictor data: application to predict optimal type 2 diabetes therapy in electronic health record data
复制标题

用于估计缺少预测数据的个体结果的狄利克雷过程混合模型:在电子健康记录数据中预测最佳 2 型糖尿病治疗的应用

DOI:
10.1101/2022.07.26.22278066
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Cardoso P
Cardoso P
中科院分区:
--
文献类型:
--
作者:
Cardoso P

文献摘要

参考文献

相似文献

背景数据缺失是回归建模中的一个常见问题。大部分的文献侧重于处理丢失的结果变量,但也有挑战时,处理丢失的预测信息,特别是当试图建立预测模型在practice.MethodsWe开发一个灵活的贝叶斯方法处理丢失的预测信息回归模型。对于预测,这为从业者提供了缺失预测因子信息和结果变量的完整后验预测分布,条件是观察到的预测因子。我们将我们的方法应用于先前提出的2型糖尿病二线治疗的治疗选择模型。我们的方法结合了回归模型和Dirichlet过程混合模型(DPMM),前者定义了治疗选择模型,后者提供了一种灵活的方法来模拟预测因子的联合分布。结果我们表明,在完全随机缺失(MCAR)和随机缺失(MAR)假设下,(相对于缺失的预测因子),DPMM可以对预测因子变量之间的复杂关系进行建模,并根据现有信息有条件地预测缺失值。我们还表明,在存在多个缺失的预测因子,DPMM模型可以用来探索哪些变量(S),如果收集,可以提供最额外的信息有关的可能outcome.ConclusionsOur的方法可以提供医生补充信息,以帮助治疗选择决策中存在的缺失数据,在临床实践中建立和实施预测模型时,缺失的预测变量是一个重大挑战。删除缺失信息的个体,并执行一个完整的案例分析可能导致不精确和偏差。多重插补方法通常通过预测模型参数标准误差来转换不确定性,而不是一致的联合概率模型。或者,使用Dirichlet过程混合模型(DPMM)的贝叶斯方法提供了一种灵活的方法来建模预测变量的复杂联合分布,可以用于估计后验(预测)分布的缺失预测,条件是观察到的预测。使用DPMM,以这种方式,允许使用贝叶斯分层框架将缺失的预测器数据周围的不确定性传播到感兴趣的预测模型。这允许使用具有不完整预测器信息的数据集(假设完全随机缺失/随机缺失)来开发预测模型。此外,即使新的个体具有不完整的预测因子信息(在相同的假设下),也可以对其进行预测。这种方法为缺失的预测变量和结果变量提供了完整的后验预测概率分布,从而可以导出广泛的概率模型输出,以支持临床决策。
BackgroundMissing data is a common problem in regression modelling. Much of the literature focuses on handling missing outcome variables, but there are also challenges when dealing with missing predictor information, particularly when trying to build prediction models for use in practice.MethodsWe develop a flexible Bayesian approach for handling missing predictor information in regression models. For prediction this provides practitioners with full posterior predictive distributions for both the missing predictor information and the outcome variable, conditional on the observed predictors. We apply our approach to a previously proposed treatment selection model for type 2 diabetes second-line therapies. Our approach combines a regression model and a Dirichlet process mixture model (DPMM), where the former defines the treatment selection model and the latter provides a flexible way to model the joint distribution of the predictors.ResultsWe show that under missing-completely-at-random (MCAR) and missing-at-random (MAR) assumptions (with respect to the missing predictors), the DPMM can model complex relationships between predictor variables, and predict missing values conditionally on existing information. We also demonstrate that in the presence of multiple missing predictors, the DPMM model can be used to explore which variable(s), if collected, could provide the most additional information about the likely outcome.ConclusionsOur approach can provide practitioners with supplementary information to aid treatment selection decisions in the presence of missing data, and can be readily extended to other types of response model.Key MessagesMissing predictor variables present a significant challenge when building and implementing prediction models in clinical practice.Removing individuals with missing information and performing a complete case analysis can lead to imprecision and bias. Multiple imputation approaches typically translate uncertainty through prediction model parameter standard errors, as opposed to a consistent joint probability model.Alternatively, a Bayesian approach using Dirichlet process mixture models (DPMMs) offers a flexible way to model complex joint distributions of predictor variables, which can be used to estimate posterior (predictive) distributions for the missing predictors, conditional on the observed predictors.Using a DPMM, in this way allows uncertainties around missing predictor data to be propagated through to a prediction model of interest using a Bayesian hierarchical framework. This allows prediction models to be developed using datasets with incomplete predictor information (assuming missing-completely-at-random/missing-at-random). Furthermore, predictions can be made on new individuals even if they have incomplete predictor information (under the same assumptions).This approach provides full posterior predictive probability distributions for both missing predictor variables and the outcome variable, allowing a wide range of probabilistic models outputs to be derived to support clinical decision making.
敏捷
DOI: --
发表时间: --
期刊: Using R for Bayesian Spatial and Spatio-Temporal Health Modeling
影响因子: --
作者:
Andrew B. Lawson
通讯作者: Andrew B. Lawson
DOI: 10.1093/biomet/asm086
发表时间: 2008-03-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
Papaspiliopoulos, Omiros;Roberts, Gareth O.
通讯作者: Roberts, Gareth O.
DOI: 10.1016/j.spl.2008.04.001
发表时间: 2008
影响因子: 0.8
作者:
S. Favaro;S. Walker
通讯作者: S. Walker
DOI: 10.1093/biostatistics/kxq013
发表时间: 2010-07-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Molitor, John;Papathomas, Michail;Richardson, Sylvia
通讯作者: Richardson, Sylvia
DOI: 10.18637/jss.v064.i07
发表时间: 2015-03-20
影响因子: 5.8
作者:
Liverani S;Hastie DI;Azizi L;Papathomas M;Richardson S
通讯作者: Richardson S