Stitching together multiple data dimensions reveals interacting metabolomic and transcriptomic networks that modulate cell regulation.

Stitching together multiple data dimensions reveals interacting metabolomic and transcriptomic networks that modulate cell regulation.
复制标题

DOI:
10.1371/journal.pbio.1001301
复制
发表时间:
2012
期刊:
影响因子:
9.8
通讯作者:
Schadt EE
Schadt EE
中科院分区:
生物学1区
文献类型:
--
作者:
Zhu J;Sova P;Xu Q;Dombek KM;Xu EY;Vu H;Tu Z;Brem RB;Bumgarner RE;Schadt EE

文献摘要

参考文献

被引文献

相似文献

DNA变异可以被用作分离群体的扰动的系统来源,作为通过整合大规模、高维分子谱数据来推断调控网络的一种方式。细胞采用多个水平的调控,包括转录和翻译调控,驱动核心生物过程,使细胞能够对遗传和环境变化做出反应。小分子代谢物是一类重要的细胞中间产物,可以影响细胞调节,也可以成为细胞调节的靶点。由于代谢物代表蛋白质介导的细胞过程的直接输出,内源性代谢物浓度可以密切反映细胞的生理状态,特别是当与其他分子分析数据相结合时。在这里,我们开发和应用的网络重建方法,同时整合了六种不同类型的数据:内源性代谢物浓度,RNA表达,DNA变异,DNA-蛋白质结合,蛋白质-代谢物相互作用,蛋白质-蛋白质相互作用的数据,构建概率因果网络,阐明细胞调控的复杂性在一个分离的酵母菌群。由于许多代谢物被发现是在强大的遗传控制下,我们能够采用因果调节检测算法来确定因果调节器的结果网络,阐明其序列的变化影响基因表达和代谢物浓度的机制。我们研究了所有四个表达数量性状位点(eQTL)的热点与共定位的代谢物QTL,其中两个概括了已知的生物过程,而其他两个阐明了新的推定的eQTL热点的生物学机制。现在可以在个体群体中以全面的方式对整个基因组的DNA变异、RNA水平和替代亚型、代谢物水平、蛋白质水平和蛋白质状态信息、蛋白质-蛋白质相互作用以及蛋白质-DNA相互作用进行评分。这些分子实体之间的相互作用定义了生物过程的复杂网络,这些生物过程引起所有更高级的表型,包括疾病。如果我们要从大规模数据中提取意义以阐明生命系统的复杂性,那么同时整合不同维度数据的分析方法的发展至关重要。在这里,我们使用一种新的贝叶斯网络重建算法,同时整合DNA变异,RNA水平,代谢物水平,蛋白质-蛋白质相互作用数据,蛋白质-DNA结合数据,蛋白质-小分子相互作用数据,在酵母中构建分子网络。我们证明,这些网络可以用来推断基因之间的因果关系,使新的基因,调节细胞调控的识别。我们表明,我们的网络预测要么概括了已知的生物学,要么可以进行前瞻性验证,表明预测网络的准确性很高。
DNA variation can be used as a systematic source of perturbation in segregating populations as a way to infer regulatory networks via the integration of large-scale, high-dimensional molecular profiling data. Cells employ multiple levels of regulation, including transcriptional and translational regulation, that drive core biological processes and enable cells to respond to genetic and environmental changes. Small-molecule metabolites are one category of critical cellular intermediates that can influence as well as be a target of cellular regulations. Because metabolites represent the direct output of protein-mediated cellular processes, endogenous metabolite concentrations can closely reflect cellular physiological states, especially when integrated with other molecular-profiling data. Here we develop and apply a network reconstruction approach that simultaneously integrates six different types of data: endogenous metabolite concentration, RNA expression, DNA variation, DNA–protein binding, protein–metabolite interaction, and protein–protein interaction data, to construct probabilistic causal networks that elucidate the complexity of cell regulation in a segregating yeast population. Because many of the metabolites are found to be under strong genetic control, we were able to employ a causal regulator detection algorithm to identify causal regulators of the resulting network that elucidated the mechanisms by which variations in their sequence affect gene expression and metabolite concentrations. We examined all four expression quantitative trait loci (eQTL) hot spots with colocalized metabolite QTLs, two of which recapitulated known biological processes, while the other two elucidated novel putative biological mechanisms for the eQTL hot spots. It is now possible to score variations in DNA across whole genomes, RNA levels and alternative isoforms, metabolite levels, protein levels and protein state information, protein–protein interactions, and protein–DNA interactions, in a comprehensive fashion in populations of individuals. Interactions among these molecular entities define the complex web of biological processes that give rise to all higher order phenotypes, including disease. The development of analytical approaches that simultaneously integrate different dimensions of data is essential if we are to extract the meaning from large-scale data to elucidate the complexity of living systems. Here, we use a novel Bayesian network reconstruction algorithm that simultaneously integrates DNA variation, RNA levels, metabolite levels, protein–protein interaction data, protein–DNA binding data, and protein–small-molecule interaction data to construct molecular networks in yeast. We demonstrate that these networks can be used to infer causal relationships among genes, enabling the identification of novel genes that modulate cellular regulation. We show that our network predictions either recapitulate known biology or can be prospectively validated, demonstrating a high degree of accuracy in the predicted network.
DOI: 10.1093/nar/gkj003
发表时间: 2006-01-01
影响因子: 14.9
作者:
Güldener U;Münsterkötter M;Oesterheld M;Pagel P;Ruepp A;Mewes HW;Stümpflen V
通讯作者: Stümpflen V
DOI: 10.1038/nprot.2008.107
发表时间: 2008-01-01
期刊: NATURE PROTOCOLS
影响因子: 14.8
作者:
Bennett, Bryson D.;Yuan, Jie;Rabinowitz, Joshua D.
通讯作者: Rabinowitz, Joshua D.
DOI: 10.1021/ac900999t
发表时间: 2009-09-01
影响因子: 7.4
作者:
Canelas, Andre B.;ten Pierick, Angela;Heijnen, Joseph J.
通讯作者: Heijnen, Joseph J.
DOI: 10.1371/journal.pgen.1000034
发表时间: 2008-03-14
期刊: PLoS genetics
影响因子: 4.5
作者:
Ferrara CT;Wang P;Neto EC;Stevens RD;Bain JR;Wenner BR;Ilkayeva OR;Keller MP;Blasiole DA;Kendziorski C;Yandell BS;Newgard CB;Attie AD
通讯作者: Attie AD
DOI: 10.1007/s00335-008-9095-z
发表时间: 2008-03-01
期刊: MAMMALIAN GENOME
影响因子: 2.5
作者:
Gordon, Ryan R.;Hunter, Kent W.;Pomp, Daniel
通讯作者: Pomp, Daniel