Assessing key decisions for transcriptomic data integration in biochemical networks

Assessing key decisions for transcriptomic data integration in biochemical networks
复制标题

DOI:
10.1371/journal.pcbi.1007185
复制
发表时间:
2019-07-01
影响因子:
4.3
通讯作者:
Lewis, Nathan E.
Lewis, Nathan E.
中科院分区:
生物学2区
文献类型:
--
作者:
Richelle, Anne;Joshi, Chintan;Lewis, Nathan E.

文献摘要

被引文献

相似文献

为了深入了解复杂的生物过程,基因组规模的数据(例如,RNA-Seq)通常覆盖在生物化学网络上。然而,由于同工酶和蛋白质复合物的存在,许多网络在基因和网络边缘之间不具有一对一的关系。因此,必须决定如何将数据覆盖到网络上。例如,对于代谢网络,这些决策包括(1)如何使用基因-蛋白质-反应规则整合基因表达水平,(2)用于选择表达数据的阈值以将相关基因视为“活性”的方法,以及(3)施加这些步骤的顺序。然而,这些决定的影响尚未得到系统的检验。我们使用跨32种组织的转录组数据集比较了20种决策组合,并表明哪种反应可以被认为是活跃的(即,在覆盖数据之后具有非零表达水平的基因组规模代谢网络的反应)主要受所使用的阈值方法影响。为了确定最合适的决策,我们评估了这些决策如何影响组织特异性活性反应列表的获取,这些列表概括了器官系统组织组。这些结果将提供指导方针,以改善数据分析与生化网络,并促进特定环境的代谢模型的建设。
To gain insights into complex biological processes, genome-scale data (e.g., RNA-Seq) are often overlaid on biochemical networks. However, many networks do not have a one-to-one relationship between genes and network edges, due to the existence of isozymes and protein complexes. Therefore, decisions must be made on how to overlay data onto networks. For example, for metabolic networks, these decisions include (1) how to integrate gene expression levels using gene-protein-reaction rules, (2) the approach used for selection of thresholds on expression data to consider the associated gene as "active", and (3) the order in which these steps are imposed. However, the influence of these decisions has not been systematically tested. We compared 20 decision combinations using a transcriptomic dataset across 32 tissues and showed that definition of which reaction may be considered as active (i.e., reactions of the genome-scale metabolic network with a non-zero expression level after overlaying the data) is mainly influenced by thresholding approach used. To determine the most appropriate decisions, we evaluated how these decisions impact the acquisition of tissue-specific active reaction lists that recapitulate organ-system tissue groups. These results will provide guidelines to improve data analyses with biochemical networks and facilitate the construction of context-specific metabolic models.