sEBM: Scaling Event Based Models to Predict Disease Progression via Implicit Biomarker Selection and Clustering.

sEBM: Scaling Event Based Models to Predict Disease Progression via Implicit Biomarker Selection and Clustering.
复制标题

sEBM:通过隐式生物标志物选择和聚类扩展基于事件的模型来预测疾病进展。

DOI:
10.1007/978-3-031-34048-2_17
复制
发表时间:
2023
期刊:
Information processing in medical imaging : proceedings of the ... conference
影响因子:
--
通讯作者:
Mitchell,CassieS
Mitchell,CassieS
中科院分区:
--
文献类型:
--
作者:
Tandon,Raghav;Kirkpatrick,Anna;Mitchell,CassieS

文献摘要

相似文献

基于事件的模型(EBM)是一种概率生成性模型,用于探索随着疾病进展而发生的生物标志物的变化。疾病的进展被认为是通过一系列生物标记物失调的“事件”发生的。循证医学估计了生物标记物失调事件序列。它计算给定失调序列的数据似然,并随后评估该失调序列的后验分布。针对后验分布的难解性,采用马尔科夫链蒙特卡罗方法生成后验分布下的样本。然而,可能的序列集合增加了ASN!其中N是生物标志物的数量(数据维度),并且迅速变得大得令人望而却步,无法通过MCMC进行有效采样。本工作提出了基于事件的大规模生物标记物模型(如高维数据)。首先,sEBM隐含地选择对疾病进展建模有用的生物标记物的子集,并仅针对该子集推断事件序列。其次,sEBM将在事件序列中具有相似位置的生物标记物聚集在一起,并且只对“簇”进行排序,每个连续的簇对应于疾病进展的下一阶段。用于构造sEBM方法的这两个修改被证明将事件序列的可能空间减少了多个数量级。这些新颖的修改得到了理论和实验的支持,在合成和真实的临床数据上为sEBM在更高维度的环境下工作提供了验证。在已知基本事实的合成数据上的结果表明,随着数据维度的增加,sEBM的性能优于以前的EBM变体。SEBM成功地实施了多达300个生物标志物,比以前的EBM应用程序增加了6倍。SEBM的实际临床应用使用了来自公开可用的阿尔茨海默病神经成像倡议(ADNI)数据的119个神经成像标记物,将受试者划分为疾病进展的6个阶段。受试者包括认知正常(CN)、轻度认知障碍(MCI)和阿尔茨海默病(AD)。按EBM分期分为3组(4.6E、−、32)。增加的sEBM阶段是MCI受试者转化为AD()的强烈预测因素,这一点得到了年龄、性别、教育程度和APOE4状态调整后的COX比例风险模型的验证。与EBM一样,sEBM不依赖于先验定义的诊断标签,只使用横断面数据。
The Event Based Model (EBM) is a probabilistic generative model to explore biomarker changes occurring as a disease progresses. Disease progression is hypothesized to occur through a sequence of biomarker dysregulation “events”. The EBM estimates the biomarker dysregulation event sequence. It computes the data likelihood for a given dysregulation sequence, and subsequently evaluates the posterior distribution on the dysregulation sequence. Since the posterior distribution is intractable, Markov Chain Monte-Carlo is employed to generate samples under the posterior distribution. However, the set of possible sequences increases asN! whereNis the number of biomarkers (data dimension) and quickly becomes prohibitively large for effective sampling via MCMC. This work proposes the “scaled EBM” (sEBM) to enable event based modeling on large biomarker sets (e.g. high-dimensional data). First, sEBM implicitly selects a subset of biomarkers useful for modeling disease progression and infers the event sequence only for that subset. Second, sEBM clusters biomarkers with similar positions in the event sequence and only orders the “clusters”, with each successive cluster corresponding to the next stage in disease progression. These two modifications used to construct the sEBM method provably reduces the possible space of event sequences by multiple orders of magnitude. The novel modifications are supported by theory and experiments on synthetic and real clinical data provides validation for sEBM to work in higher dimensional settings. Results on synthetic data with known ground truth shows that sEBM outperforms previous EBM variants as data dimensions increase. sEBM was successfully implemented with up to 300 biomarkers, which is a 6-fold increase over previous EBM applications. A real-world clinical application of sEBM is performed using 119 neuroimaging markers from publicly available Alzheimer’s Disease Neuroimaging Initiative (ADNI) data to stratify subjects into 6 stages of disease progression. Subjects included cognitively normal (CN), mild cognitive impairment (MCI), and Alzheimer’s Disease (AD). sEBM stage is differentiated for the 3 groups (4.6e−32). Increased sEBM stage is a strong predictor of conversion risk to AD () for MCI subjects, as verified with a Cox proportional-hazards model adjusted for age, sex, education and APOE4 status. Like EBM, sEBM does not rely on apriori defined diagnostic labels and only uses cross-sectional data.