Multi-omics Approach for Estimating Metabolic Networks Using Low-Order Partial Correlations

Multi-omics Approach for Estimating Metabolic Networks Using Low-Order Partial Correlations
复制标题

DOI:
10.1089/cmb.2013.0043
复制
发表时间:
2013-08-01
影响因子:
1.7
通讯作者:
Miyano, Satoru
Miyano, Satoru
中科院分区:
生物学4区
文献类型:
--
作者:
Kayano, Mitsunori;Imoto, Seiya;Miyano, Satoru

文献摘要

被引文献

相似文献

代谢组分析的两个典型目的是估计代谢途径和了解新陈代谢背后的调节系统。这些分析的一个强大的信息来源是一组关于RNA、蛋白质和代谢物的多组学数据。然而,同时分析多个组学数据并揭示代谢背后的系统的综合方法还没有很好地建立起来。我们开发了一种基于低阶部分相关性和稳健相关系数的统计方法,用于从代谢组、蛋白质组和转录组数据中估计代谢网络。我们的方法是用低阶,特别是一阶部分相关性的最大值(MF-PCOR)来定义,以便分配具有最高相关性的正确边缘,并检测对相关系数有强烈影响的因素。首先,通过与真实和合成数据的数值实验,我们表明,使用蛋白质和酶的转录数据提高了在MF-PCOR中估计代谢网络的准确性。在这些实验中,通过与相关网络(COR)和高斯图模型(GGM)的比较,也证明了该方法的有效性。我们的理论研究证实,MF-PCOR的性能可能优于其他竞争方法。此外,在实际数据分析中,我们调查了代谢产物、酶和酶基因在MF-PCOR建立的网络中被确定为重要因素的作用。然后我们发现,它们中的一些对应于由催化酶介导的代谢物之间的特定反应,这些反应很难通过仅基于代谢物数据的分析来识别。
Two typical purposes of metabolome analysis are to estimate metabolic pathways and to understand the regulatory systems underlying the metabolism. A powerful source of information for these analyses is a set of multi-omics data for RNA, proteins, and metabolites. However, integrated methods that analyze multi-omics data simultaneously and unravel the systems behind metabolisms have not been well established. We developed a statistical method based on low-order partial correlations with a robust correlation coefficient for estimating metabolic networks from metabolome, proteome, and transcriptome data. Our method is defined by the maximum of low-order, particularly first-order, partial correlations (MF-PCor) in order to assign a correct edge with the highest correlation and to detect the factors that strongly affect the correlation coefficient. First, through numerical experiments with real and synthetic data, we showed that the use of protein and transcript data of enzymes improved the accuracy of the estimated metabolic networks in MF-PCor. In these experiments, the effectiveness of the proposed method was also demonstrated by comparison with a correlation network (Cor) and a Gaussian graphical model (GGM). Our theoretical investigation confirmed that the performance of MF-PCor could be superior to that of the competing methods. In addition, in the real data analysis, we investigated the role of metabolites, enzymes, and enzyme genes that were identified as important factors in the network established by MF-PCor. We then found that some of them corresponded to specific reactions between metabolites mediated by catalytic enzymes that were difficult to be identified by analysis based on metabolite data alone.