A generative-discriminative framework that integrates imaging, genetic, and diagnosis into coupled low dimensional space

A generative-discriminative framework that integrates imaging, genetic, and diagnosis into coupled low dimensional space
复制标题

DOI:
10.1016/j.neuroimage.2021.118200
复制
发表时间:
2021-06
期刊:
影响因子:
5.7
通讯作者:
Sayan Ghosal;Qiang Chen;G. Pergola;A. Goldman;William Ulrich;K. Berman;G. Blasi;L. Fazio;A. Rampino;A. Bertolino;D. Weinberger;V. Mattay;A. Venkataraman
Sayan Ghosal;Qiang Chen;G. Pergola;A. Goldman;William Ulrich;K. Berman;G. Blasi;L. Fazio;A. Rampino;A. Bertolino;D. Weinberger;V. Mattay;A. Venkataraman
中科院分区:
医学1区
文献类型:
--
作者:
Sayan Ghosal;Qiang Chen;G. Pergola;A. Goldman;William Ulrich;K. Berman;G. Blasi;L. Fazio;A. Rampino;A. Bertolino;D. Weinberger;V. Mattay;A. Venkataraman

文献摘要

相似文献

我们提出了一个新的优化框架,集成了成像和遗传学数据,同时进行生物标志物鉴定和疾病分类。我们模型的生成组件使用字典学习框架将成像和遗传数据投影到共享的低维空间中。我们通过将线性投影系数绑定到相同的潜在空间来耦合两种数据模式。我们模型的判别成分对疾病诊断的投影向量使用逻辑回归。这项预测任务隐含地指导我们的框架,以找到健康人群和疾病人群之间存在本质差异的可解释生物标志物。我们通过在联合目标函数中加入图正则化惩罚来利用不同大脑区域的互联性。我们还使用组稀疏性惩罚来找到一组具有代表性的遗传基向量,这些向量跨越低维空间,其中受试者在患者和对照组之间易于分离。我们在一项精神分裂症人群研究中评估了我们的模型,该模型包括两个任务fMRI范式和单核苷酸多态性(SNP)数据。通过十倍交叉验证,我们将我们的生成判别框架与影像学和遗传学数据的典型相关分析(CCA)、影像学和遗传学数据的平行独立成分分析(pICA)、随机森林(RF)分类和线性支持向量机(SVM)进行了比较。我们还通过亚采样量化了成像和遗传生物标志物的可重复性。我们的框架实现了更高的预测精度和识别稳健的生物标志物。此外,相关的大脑区域和遗传变异是精神分裂症缺陷的基础。
We propose a novel optimization framework that integrates imaging and genetics data for simultaneous biomarker identification and disease classification. The generative component of our model uses a dictionary learning framework to project the imaging and genetic data into a shared low dimensional space. We have coupled both the data modalities by tying the linear projection coefficients to the same latent space. The discriminative component of our model uses logistic regression on the projection vectors for disease diagnosis. This prediction task implicitly guides our framework to find interpretable biomarkers that are substantially different between a healthy and disease population. We exploit the interconnectedness of different brain regions by incorporating a graph regularization penalty into the joint objective function. We also use a group sparsity penalty to find a representative set of genetic basis vectors that span a low dimensional space where subjects are easily separable between patients and controls. We have evaluated our model on a population study of schizophrenia that includes two task fMRI paradigms and single nucleotide polymorphism (SNP) data. Using ten-fold cross validation, we compare our generative-discriminative framework with canonical correlation analysis (CCA) of imaging and genetics data, parallel independent component analysis (pICA) of imaging and genetics data, random forest (RF) classification, and a linear support vector machine (SVM). We also quantify the reproducibility of the imaging and genetics biomarkers via subsampling. Our framework achieves higher class prediction accuracy and identifies robust biomarkers. Moreover, the implicated brain regions and genetic variants underlie the well documented deficits in schizophrenia.