A Novel Method for Cancer Subtyping and Risk Prediction Using Consensus Factor Analysis

A Novel Method for Cancer Subtyping and Risk Prediction Using Consensus Factor Analysis
复制标题

DOI:
10.3389/fonc.2020.01052
复制
发表时间:
2020-06-24
影响因子:
4.7
通讯作者:
Tin Nguyen
Tin Nguyen
中科院分区:
医学3区
文献类型:
--
作者:
Duc Tran;Hung Nguyen;Tin Nguyen

文献摘要

被引文献

相似文献

癌症是一个涵盖性术语,包括一系列疾病,从快速生长和致命的疾病到进展至死亡的可能性较低或延迟的惰性病变。一项尚未解决的关键挑战是,以相关临床差异(例如生存率)为特征的分子疾病亚型很难区分。随着多组学技术的进步,亚型分型方法已转向数据集成,以便从多个层面考虑现象的整体角度区分亚型。然而,这些综合方法仍然受到统计假设和对噪声敏感性的限制。此外,他们无法使用多组学数据预测患者的风险评分。在这里,我们提出了一种名为通过共识因子分析进行亚型分型(SCFA)的新方法,该方法可以有效地从一致的分子模式中去除噪声信号,以便可靠地识别癌症亚型并准确预测患者的风险评分。在对癌症基因组图谱 (TCGA) 中提供的 30 种癌症相关的 7,973 个样本进行的广泛分析中,我们证明 SCFA 在发现具有显着不同生存状况的新亚型方面优于最先进的方法。我们还证明 SCFA 能够预测与患者真实生存率和生命状态高度相关的风险评分。更重要的是,当更多数据类型集成到分析中时,子类型发现和风险预测的准确性就会提高。 SCFA 软件和 TCGA 数据包将在 Bioconductor 上提供。
Cancer is an umbrella term that includes a range of disorders, from those that are fast-growing and lethal to indolent lesions with low or delayed potential for progression to death. One critical unmet challenge is that molecular disease subtypes characterized by relevant clinical differences, such as survival, are difficult to differentiate. With the advancement of multi-omics technologies, subtyping methods have shifted toward data integration in order to differentiate among subtypes from a holistic perspective that takes into consideration phenomena at multiple levels. However, these integrative methods are still limited by their statistical assumption and their sensitivity to noise. In addition, they are unable to predict the risk scores of patients using multi-omics data. Here, we present a novel approach named Subtyping via Consensus Factor Analysis (SCFA) that can efficiently remove noisy signals from consistent molecular patterns in order to reliably identify cancer subtypes and accurately predict risk scores of patients. In an extensive analysis of 7,973 samples related to 30 cancers that are available at The Cancer Genome Atlas (TCGA), we demonstrate that SCFA outperforms state-of-the-art approaches in discovering novel subtypes with significantly different survival profiles. We also demonstrate that SCFA is able to predict risk scores that are highly correlated with true patient survival and vital status. More importantly, the accuracy of subtype discovery and risk prediction improves when more data types are integrated into the analysis. The SCFA software and TCGA data packages will be available on Bioconductor.