Theory and limitations of genetic network inference from microarray data

Theory and limitations of genetic network inference from microarray data
复制标题

DOI:
10.1196/annals.1407.019
复制
发表时间:
2007-01-01
期刊:
REVERSE ENGINEERING BIOLOGICAL NETWORKS
影响因子:
--
通讯作者:
Califano, Andrea
Califano, Andrea
中科院分区:
其他
文献类型:
--
作者:
Margolin, Adam A.;Califano, Andrea

文献摘要

被引文献

相似文献

自从10多年前基因表达微阵列技术出现以来,已经开发了许多计算方法,旨在利用mRNA丰度谱之间的统计关联来预测转录调控相互作用。最终目标是开发因果网络模型,描述基因(通过它们的蛋白质产物)对彼此施加的转录影响,这可用于预测网络中断(例如,突变)导致疾病表型,以及适当的治疗干预。然而,微阵列数据仅测量基因调控网络中相互作用变量的一小部分,因为已知细胞通过许多不同的机制调控基因表达。尽管许多研究人员已经承认mRNA谱之间的统计依赖性的解释是有问题的,但很少有工作在理论上描述使用模型解释未观察到的相互作用变量的推断依赖性的性质。在这项工作中,我们回顾了逆向工程算法背后的理论来自三个独立的学科系统控制理论,图形模型和信息理论,并强调各种方法之间的数学关系。然后,我们应用最近的理论工作,构建图形模型的背景下,反向工程遗传网络的潜在变量。我们证明,即使添加简单的潜变量也会导致非直接相互作用(例如,共调节的)基因,这些基因不能通过调节任何观察到的变量来消除。
Since the advent of gene expression microarray technology more than 10 years ago, many computational approaches have been developed aimed at using statistical associations between mRNA abundance profiles to predict transcriptional regulatory interactions. The ultimate goal is to develop causal network models describing the transcriptional influences that genes exert on each other (via their protein products), which can be used to predict network disruptions (e.g., mutations) leading to a disease phenotype, as well as the appropriate therapeutic intervention. However, microarray data measure only a small component of the interacting variables in a genetic regulatory network, as cells are known to regulate gene expression via many diverse mechanisms. Although many researchers have acknowledged the questionable interpretation of statistical dependencies between mRNA profiles, very little work has been done on theoretically characterizing the nature of inferred dependencies using models that account for unobserved interacting variables. In this work, we review the theory behind reverse engineering algorithms derived from three separate disciplines-system control theory, graphical models, and information theory-and highlight several mathematical relationships between the various methods. We then apply recent theoretical work on constructing graphical models with latent variables to the context of reverse engineering genetic networks. We demonstrate that even the addition of simple latent variables induces statistical dependencies between non-directly interacting (e.g., co-regulated) genes that cannot be eliminated by conditioning on any observed variables.