Pattern fusion analysis by adaptive alignment of multiple heterogeneous omics data

Pattern fusion analysis by adaptive alignment of multiple heterogeneous omics data
复制标题

通过多个异构组学数据的自适应对齐进行模式融合分析

DOI:
10.1093/bioinformatics/btx176
复制
发表时间:
2017-09-01
期刊:
影响因子:
5.8
通讯作者:
Chen, Luonan
Chen, Luonan
中科院分区:
生物学3区
文献类型:
--
作者:
Shi, Qianqian;Zhang, Chuanchao;Chen, Luonan

文献摘要

被引文献

相似文献

动机:整合不同的组学图谱是一项具有挑战性的任务,它以多视角的方式为理解复杂疾病提供了一种全面的途径。这种整合的一个关键是根据数据结构提取内在模式,以便在各种数据类型中发现一致的信息,即使存在噪声干扰。因此,我们提出了一种名为“模式融合分析”(PFA)的新框架,它执行自动的信息比对和偏差校正,将局部样本模式(例如来自每种数据类型)融合为与表型相对应的全局样本模式(例如跨越大多数数据类型)。特别是,PFA能够通过对模式最优地调整每种数据类型的影响,从不同的组学图谱中识别出显著的样本模式,从而缓解处理不同平台以及异构数据不同可靠性水平的问题。 结果:为了验证我们方法的有效性,我们首先在各种合成数据集上对PFA进行了测试,发现与诸如iClusterPlus、SNF和moCluster等最先进的方法相比,PFA不仅能够从多组学数据中捕捉内在的样本聚类结构,还能提供一种自动的权重方案来衡量数据类型甚至样本的相应贡献。此外,计算结果表明,在癌细胞系百科全书(CCLE)数据集中,PFA能够揭示具有不同信噪比的数据类型之间共享和互补的样本模式,并且在癌症基因组图谱(TCGA)数据集中识别临床不同的癌症亚型方面优于其他研究。
Motivation: Integrating different omics profiles is a challenging task, which provides a comprehensive way to understand complex diseases in a multi-view manner. One key for such an integration is to extract intrinsic patterns in concordance with data structures, so as to discover consistent information across various data types even with noise pollution. Thus, we proposed a novel framework called 'pattern fusion analysis' (PFA), which performs automated information alignment and bias correction, to fuse local sample-patterns (e.g. from each data type) into a global sample-pattern corresponding to phenotypes (e.g. across most data types). In particular, PFA can identify significant sample-patterns from different omics profiles by optimally adjusting the effects of each data type to the patterns, thereby alleviating the problems to process different platforms and different reliability levels of heterogeneous data.Results: To validate the effectiveness of our method, we first tested PFA on various synthetic data-sets, and found that PFA can not only capture the intrinsic sample clustering structures from the multi-omics data in contrast to the state-of-the-art methods, such as iClusterPlus, SNF and moCluster, but also provide an automatic weight-scheme to measure the corresponding contributions by data types or even samples. In addition, the computational results show that PFA can reveal shared and complementary sample-patterns across data types with distinct signal-to-noise ratios in Cancer Cell Line Encyclopedia (CCLE) datasets, and outperforms over other works at identifying clinically distinct cancer subtypes in The Cancer Genome Atlas (TCGA) datasets.