Causal reasoning on biological networks: interpreting transcriptional changes

Causal reasoning on biological networks: interpreting transcriptional changes
复制标题

DOI:
10.1093/bioinformatics/bts090
复制
发表时间:
2012-04-15
期刊:
影响因子:
5.8
通讯作者:
Huang, Enoch S.
Huang, Enoch S.
中科院分区:
生物学3区
文献类型:
--
作者:
Chindelevitch, Leonid;Ziemek, Daniel;Huang, Enoch S.

文献摘要

被引文献

相似文献

动机:在过去的十年中,高通量数据集的解释仍然是计算生物学的核心挑战之一。此外,随着生物学知识量的增加,以有意义的方式整合这一庞大的知识体系变得越来越困难。在这篇文章中,我们提出了一个特定的解决方案,这两个challenges.Methods:我们整合现有的生物学知识,通过构建一个网络的分子相互作用的一种特定的:因果关系的相互作用。由此产生的因果关系图可以查询,以提出解释在高通量基因表达实验中观察到的变化的分子假设。我们发现,一个简单的评分函数可以区分大量的竞争分子的假设上游的基因表达谱中观察到的变化的原因。然后,我们开发了一种分析方法来计算每个分数的统计显著性。这种分析方法还有助于评估随机或对抗性噪声对我们的模型的预测能力的影响。结果:我们的研究结果表明,我们从已知的生物文献构建的因果图是非常强大的随机噪声和丢失或虚假的信息。我们展示了我们的因果推理模型的两个具体的例子,一个来自癌症数据集,另一个来自心脏肥大实验的力量。我们的结论是,因果推理模型提供了一个有价值的除了生物学家的工具包的基因表达数据的解释。
Motivation: The interpretation of high-throughput datasets has remained one of the central challenges of computational biology over the past decade. Furthermore, as the amount of biological knowledge increases, it becomes more and more difficult to integrate this large body of knowledge in a meaningful manner. In this article, we propose a particular solution to both of these challenges.Methods: We integrate available biological knowledge by constructing a network of molecular interactions of a specific kind: causal interactions. The resulting causal graph can be queried to suggest molecular hypotheses that explain the variations observed in a high-throughput gene expression experiment. We show that a simple scoring function can discriminate between a large number of competing molecular hypotheses about the upstream cause of the changes observed in a gene expression profile. We then develop an analytical method for computing the statistical significance of each score. This analytical method also helps assess the effects of random or adversarial noise on the predictive power of our model.Results: Our results show that the causal graph we constructed from known biological literature is extremely robust to random noise and to missing or spurious information. We demonstrate the power of our causal reasoning model on two specific examples, one from a cancer dataset and the other from a cardiac hypertrophy experiment. We conclude that causal reasoning models provide a valuable addition to the biologist's toolkit for the interpretation of gene expression data.