Network-based pathway enrichment analysis with incomplete network information

Network-based pathway enrichment analysis with incomplete network information
复制标题

DOI:
10.1093/bioinformatics/btw410
复制
发表时间:
2016-10-15
期刊:
影响因子:
5.8
通讯作者:
Michailidis, George
Michailidis, George
中科院分区:
生物学3区
文献类型:
--
作者:
Ma, Jing;Shojaie, Ali;Michailidis, George

文献摘要

被引文献

相似文献

动机:途径富集分析已成为生物医学研究人员深入了解差异表达基因、蛋白质和代谢产物的潜在生物学的关键工具。它降低了复杂性,并提供了响应治疗和/或疾病状态的细胞活性变化的系统级视图。使用现有途径网络信息的方法已被证明优于仅考虑途径成员的更简单的方法。然而,尽管在理解生物途径成员之间的关联方面取得了重大进展,并且包含有关生物分子相互作用信息的数据库得到了扩展,但现有的网络信息可能不完整或不准确,并且不是细胞类型或疾病状况特异性的。我们提出了一个约束网络估计框架,结合网络估计的基础上细胞和条件特定的高,三维组学数据与来自现有数据库的交互信息。随后使用所得到的途径拓扑信息来提供用于同时测试途径成员的表达水平差异以及它们的相互作用的框架。我们研究了所提出的网络估计的渐近性质和路径富集的测试,并研究其在模拟和真实的数据设置的小样本性能。
Motivation: Pathway enrichment analysis has become a key tool for biomedical researchers to gain insight into the underlying biology of differentially expressed genes, proteins and metabolites. It reduces complexity and provides a system-level view of changes in cellular activity in response to treatments and/or in disease states. Methods that use existing pathway network information have been shown to outperform simpler methods that only take into account pathway membership. However, despite significant progress in understanding the association amongst members of biological pathways, and expansion of data bases containing information about interactions of bio-molecules, the existing network information may be incomplete or inaccurate and is not cell-type or disease condition-specific.Results: We propose a constrained network estimation framework that combines network estimation based on cell- and condition-specific high-dimensional Omics data with interaction information from existing data bases. The resulting pathway topology information is subsequently used to provide a framework for simultaneous testing of differences in expression levels of pathway members, as well as their interactions. We study the asymptotic properties of the proposed network estimator and the test for pathway enrichment, and investigate its small sample performance in simulated and real data settings.