A probabilistic methodology for integrating knowledge and experiments on biological networks

A probabilistic methodology for integrating knowledge and experiments on biological networks
复制标题

DOI:
10.1089/cmb.2006.13.165
复制
发表时间:
2006-03-01
影响因子:
1.7
通讯作者:
Shamir, R
Shamir, R
中科院分区:
生物学4区
文献类型:
--
作者:
Gat-Viks, I;Tanay, A;Shamir, R

文献摘要

被引文献

相似文献

传统上,研究生物系统的方法是聚焦于一个特定的子系统,为其建立一个直观的模型,并使用精心设计的实验结果来改进模型。现代实验技术提供了关于生物系统全球行为的大量数据,系统地使用这些大型数据集来提炼现有知识是一个重大挑战。在这里,我们介绍了一个扩展的计算框架,它结合了现有定性模型的形式化、概率建模和高通量实验数据的集成。使用我们的方法,可以在关于系统的先验知识的背景下解释全基因组测量,为这些知识的准确性赋予统计意义,并学习改进后的符合实验的改进模型。我们的模型被表示为一个概率因子图,并且该框架容纳了对不同生物元素的部分测量。我们研究了几种概率推理算法的性能,表明即使在存在反馈回路和复杂逻辑的情况下,也可以可靠地推断隐藏的模型变量。我们展示了如何使用假设检验来提炼关于组合调控关系的先验知识,并为学习的模型特征得出p值。我们在一个模拟模型和两个真实的酵母模型上测试了我们的方法和算法。特别是,我们使用我们的方法来探索酵母对高渗休克反应和酵母赖氨酸生物合成系统中调节因子之间的未知关系。我们对生物调控分析的综合方法被证明是将定性和定量证据协同地结合到具体的生物预测中。
Biological systems are traditionally studied by focusing on a specific subsystem, building an intuitive model for it, and refining the model using results from carefully designed experiments. Modern experimental techniques provide massive data on the global behavior of biological systems, and systematically using these large datasets for refining existing knowledge is a major challenge. Here we introduce an extended computational framework that combines formalization of existing qualitative models, probabilistic modeling, and integration of high-throughput experimental data. Using our methods, it is possible to interpret genomewide measurements in the context of prior knowledge on the system, to assign statistical meaning to the accuracy of such knowledge, and to learn refined models with improved fit to the experiments. Our model is represented as a probabilistic factor graph, and the framework accommodates partial measurements of diverse biological elements. We study the performance of several probabilistic inference algorithms and show that hidden model variables can be reliably inferred even in the presence of feedback loops and complex logic. We show how to refine prior knowledge on combinatorial regulatory relations using hypothesis testing and derive p-values for learned model features. We test our methodology and algorithms on a simulated model and on two real yeast models. In particular, we use our method to explore uncharacterized relations among regulators in the yeast response to hyper-osmotic shock and in the yeast lysine biosynthesis system. Our integrative approach to the analysis of biological regulation is demonstrated to synergistically combine qualitative and quantitative evidence into concrete biological predictions.