Growing Bayesian network models of gene networks from seed genes

Growing Bayesian network models of gene networks from seed genes
复制标题

DOI:
10.1093/bioinformatics/bti1137
复制
发表时间:
2005-09-01
期刊:
影响因子:
5.8
通讯作者:
Tegnér, J
Tegnér, J
中科院分区:
生物学3区
文献类型:
--
作者:
Peña, JM;Björkegren, J;Tegnér, J

文献摘要

被引文献

相似文献

动机:在过去的几年里,贝叶斯网络(BN)作为基因网络的模型受到了计算生物学界越来越多的关注,尽管从基因表达数据中学习它们是有问题的。大多数基因表达数据库包含数千个基因的测量值,但现有的从数据中学习BN的算法不能扩展到这样的高维数据库。这意味着用户必须提前决定哪些基因被包括在学习过程中,通常不超过几百个,哪些基因被排除在外。我们提出了一种替代方法来克服这个问题。结果:我们提出了一种新的算法,从基因表达数据学习BN模型的基因网络。我们的算法从用户那里接收种子基因S和正整数R,并为依赖于S的基因返回BN,使得小于R的其他基因介导依赖性。我们的算法通过重复以下步骤R + 1次,然后修剪一些基因来增长BN,BN最初只包含S;找到BN中所有基因的父母和孩子,并将它们添加到其中。直观地说,我们的算法为用户提供了一个半径为R的窗口,可以查看基因网络的BN模型,而无需事先排除任何基因。我们证明了我们的算法是正确的忠实性假设下。我们评估我们的算法模拟和生物数据(罗塞塔纲要)与令人满意的结果。
Motivation: For the last few years, Bayesian networks (BNs) have received increasing attention from the computational biology community as models of gene networks, though learning them from gene-expression data is problematic. Most gene-expression databases contain measurements for thousands of genes, but the existing algorithms for learning BNs from data do not scale to such high-dimensional databases. This means that the user has to decide in advance which genes are included in the learning process, typically no more than a few hundreds, and which genes are excluded from it. This is not a trivial decision. We propose an alternative approach to overcome this problem.Results: We propose a new algorithm for learning BN models of gene networks from gene-expression data. Our algorithm receives a seed gene S and a positive integer R from the user, and returns a BN for the genes that depend on S such that less than R other genes mediate the dependency. Our algorithm grows the BN, which initially only contains S, by repeating the following step R + 1 times and, then, pruning some genes; find the parents and children of all the genes in the BN and add them to it. Intuitively, our algorithm provides the user with a window of radius R around S to look at the BN model of a gene network without having to exclude any gene in advance. We prove that our algorithm is correct under the faithfulness assumption. We evaluate our algorithm on simulated and biological data (Rosetta compendium) with satisfactory results.