sJIVE: Supervised joint and individual variation explained

sJIVE: Supervised joint and individual variation explained
复制标题

DOI:
10.1016/j.csda.2022.107547
复制
发表时间:
2022-11-01
影响因子:
1.8
通讯作者:
Lock, Eric F.
Lock, Eric F.
中科院分区:
数学3区
文献类型:
--
作者:
Palzer, Elise F.;Wendt, Christine H.;Lock, Eric F.

文献摘要

被引文献

相似文献

在分子生物医学研究中,分析多源数据,即对同一主题数据的多种观点,已经变得越来越普遍。最近的方法试图揭示数据源内部和/或数据源之间的潜在结构和关系,而其他方法则试图为使用所有数据源的结果建立预测模型。然而,既能做到这两点的现有方法目前受到限制,因为它们要么(1)只考虑所有数据集共享的数据结构,而忽略每个源的独特结构,要么(2)首先提取底层结构,而不考虑结果。所提出的方法,监督关节和个体变异解释(sJIVE),可以同时(1)识别共享(关节)和源特定(个体)底层结构,(2)使用这些结构为结果建立线性预测模型。对这两个组成部分进行加权,以在解释多源数据和结果中的变化之间达成妥协。仿真结果表明,当多源数据中存在大量噪声时,sJIVE优于现有方法。COPDGene研究数据的应用探讨了与肺功能相关的基因表达和蛋白质组学模式。(C) 2022 Elsevier B.V.版权所有
Analyzing multi-source data, which are multiple views of data on the same subjects, has become increasingly common in molecular biomedical research. Recent methods have sought to uncover underlying structure and relationships within and/or between the data sources, and other methods have sought to build a predictive model for an outcome using all sources. However, existing methods that do both are presently limited because they either (1) only consider data structure shared by all datasets while ignoring structures unique to each source, or (2) they extract underlying structures first without consideration to the outcome. The proposed method, supervised joint and individual variation explained (sJIVE), can simultaneously (1) identify shared (joint) and source specific (individual) underlying structure and (2) build a linear prediction model for an outcome using these structures. These two components are weighted to compromise between explaining variation in the multi-source data and in the outcome. Simulations show sJIVE to outperform existing methods when large amounts of noise are present in the multi-source data. An application to data from the COPDGene study explores gene expression and proteomic patterns associated with lung function. (C) 2022 Elsevier B.V. All rights reserved.