Sparse Overlapping Group Lasso for Integrative Multi-Omics Analysis

Sparse Overlapping Group Lasso for Integrative Multi-Omics Analysis
复制标题

DOI:
10.1089/cmb.2014.0197
复制
发表时间:
2015-02
期刊:
Journal of computational biology : a journal of computational molecular cell biology
影响因子:
--
通讯作者:
Heewon Park;A. Niida;S. Miyano;S. Imoto
Heewon Park;A. Niida;S. Miyano;S. Imoto
中科院分区:
其他
文献类型:
--
作者:
Heewon Park;A. Niida;S. Miyano;S. Imoto

文献摘要

被引文献

相似文献

基因网络和图表是理解癌症异质系统的重要工具,因为癌症是一种不涉及单个基因而是与致癌过程相关的基因组合的疾病。通过基因网络进行基因组数据分析的目标是识别基因网络和所选网络中的单个基因。然而,现有方法仅执行网络选择,因此所选网络中的所有基因都包括在模型中。这导致在揭示驱动基因时过度拟合,并且结果在生物学上不可解释。为了实现“组间稀疏性”和“组内稀疏性”以基于生物学知识(即,预定义的重叠组的功能),我们提出了一个稀疏的重叠组套索通过重复的预测在扩展空间。所提出的方法有效地识别驱动基因和它们的相互作用,使用已知的生物途径信息。蒙特卡罗模拟和癌症基因组图谱(TCGA)项目数据分析表明,所提出的方法对于拟合回归模型(即,特征选择和预测准确性)。在TCGA数据分析中,我们通过多组学数据构建的表达模块和基因网络发现了潜在的癌症驱动基因,并确定未被发现的基因具有作为癌症驱动基因的强有力证据。该方法是一个有用的工具,用于识别癌症驱动基因和整合的多组学分析。
Gene networks and graphs are crucial tools for understanding a heterogeneous system of cancer, since cancer is a disease that does not involve individual genes but combinations of genes associated with oncogenic process. A goal of genomic data analysis via gene networks is to identify both gene networks and individual genes within the selected networks. Existing methods, however, perform only network selection, and thus all genes in selected networks are included in models. This leads to overfitting when uncovering driver genes, and the results are not biologically interpretable. To accomplish both "groupwise sparsity" and "within group sparsity" for identifying driver genes based on biological knowledge (i.e., predefined overlapping groups of features), we propose a sparse overlapping group lasso via duplicated predictors in extended space. The proposed method effectively identifies driver genes and their interactions using known biological pathway information. Monte Carlo simulations and The Cancer Genome Atlas (TCGA) project data analysis indicate that the proposed method is effective for fitting a regression model (i.e., feature selection and prediction accuracy) constructed with duplicated predictors in overlapping groups. In the TCGA data analysis, we uncover potential cancer driver genes via expression modules and gene networks constructed by multi-omics data and identify that the uncovered genes have strong evidences as a cancer driver gene. The proposed method is a useful tool for identifying cancer driver genes and for integrative multi-omics analysis.