PhyloGibbs: a Gibbs sampling motif finder that incorporates phylogeny.

PhyloGibbs: a Gibbs sampling motif finder that incorporates phylogeny.
复制标题

DOI:
10.1371/journal.pcbi.0010067
复制
发表时间:
2005-12
影响因子:
4.3
通讯作者:
van Nimwegen, Erik
van Nimwegen, Erik
中科院分区:
生物学2区
文献类型:
--
作者:
Siddharthan, Rahul;Siggia, Eric D;van Nimwegen, Erik

文献摘要

被引文献

相似文献

基因调控生物信息学的一个核心问题是找到调控蛋白的结合位点。鉴别这些短而模糊的序列模式最有希望的方法之一是对近缘种的同源基因间区进行比较分析。这一分析因各种因素而复杂化。首先,人们需要考虑物种之间的系统发育关系,以便区分由于功能位点的发生而产生的保护和由于进化接近而产生的虚假保护。其次,人们必须处理同源基因间区域多重比对的复杂性,并且必须考虑功能位点可能出现在保守片段之外的可能性。在这里,我们提出了一个新的基序采样算法,PhyloGibbs,运行在任意集合的多个局部序列比对的同源序列。该算法搜索了任意数量的转录因子(tf)的任意数量的结合位点可以分配给多个序列比对的所有方法。这些结合位点配置通过贝叶斯概率模型进行评分,该模型通过结合位点和“背景”基因间DNA的进化模型处理对齐序列。该模型明确考虑了物种间的系统发育关系。该算法采用模拟退火和蒙特卡罗马尔可夫链抽样对报告的所有结合位点严格分配后验概率。在对五种Saccharomyces物种的合成数据和真实数据的测试中,我们的算法明显优于其他四种基序查找算法,包括考虑系统发育的算法。我们的结果还表明,与其他算法相比,PhyloGibbs可以对其预测的可靠性做出现实的估计。我们的测试表明,在单个基因上游区域的五种多重比对中,PhyloGibbs平均恢复了酿酒酵母中50%以上的所有结合位点,特异性约为50%,恢复了33%的所有结合位点,特异性约为85%。我们还在最近基于ChIP-on-chip数据注释的多个基因间区域比对集合上测试了PhyloGibbs,以包含相同TF的结合位点。我们将PhyloGibbs的结果与之前使用其他六种基序查找算法对这些数据的分析进行了比较。对于所有其他基序寻找方法都未能找到重要基序的21个tf中的16个,PhyloGibbs确实恢复了符合文献共识的基序。在11个结果不一致的案例中,我们从文献中收集了已知靶基因的列表,发现在它们的调控区域运行PhyloGibbs产生了一个与文献一致的结合基序,只有一个例外。有趣的是,这些文献基因列表与基于ChIP-on-chip数据注释的靶标几乎没有重叠。PhyloGibbs代码可以从http://www.biozentrum.unibas.ch/~nimwegen/cgi-bin/phylogibbs.cgi或http://www.imsc.res.in/~rsidd/phylogibbs下载。我们在酵母上测试的全部预测位点可以在http://www.swissregulon.unibas.ch上找到。基因间DNA调控位点的计算发现是生物信息学的核心问题之一。直到最近,motif发现者通常采用以下两种一般方法之一。给定一组已知的共调控基因,人们搜索它们的启动子区域,寻找显著过度代表的序列基序。另外,在“系统发育足迹”方法中,人们搜索同源基因间区域的多个比对,以寻找比基于物种系统发育的预期更为保守的短片段。在这项工作中,作者提出了一种算法,PhyloGibbs,将这两种方法结合到一个集成的贝叶斯框架中。该算法在考虑序列之间的系统发育关系的同时,搜索所有可以将任意数量的转录因子的任意数量的结合位点分配给任意多序列比对集合的方法。作者对来自酵母菌基因组的合成数据和真实数据进行了大量测试,其中PhyloGibbs显著优于其他现有方法。最后,一种新颖的退火跟踪策略允许PhyloGibbs对其预测的可靠性做出准确的估计。
A central problem in the bioinformatics of gene regulation is to find the binding sites for regulatory proteins. One of the most promising approaches toward identifying these short and fuzzy sequence patterns is the comparative analysis of orthologous intergenic regions of related species. This analysis is complicated by various factors. First, one needs to take the phylogenetic relationship between the species into account in order to distinguish conservation that is due to the occurrence of functional sites from spurious conservation that is due to evolutionary proximity. Second, one has to deal with the complexities of multiple alignments of orthologous intergenic regions, and one has to consider the possibility that functional sites may occur outside of conserved segments. Here we present a new motif sampling algorithm, PhyloGibbs, that runs on arbitrary collections of multiple local sequence alignments of orthologous sequences. The algorithm searches over all ways in which an arbitrary number of binding sites for an arbitrary number of transcription factors (TFs) can be assigned to the multiple sequence alignments. These binding site configurations are scored by a Bayesian probabilistic model that treats aligned sequences by a model for the evolution of binding sites and “background” intergenic DNA. This model takes the phylogenetic relationship between the species in the alignment explicitly into account. The algorithm uses simulated annealing and Monte Carlo Markov-chain sampling to rigorously assign posterior probabilities to all the binding sites that it reports. In tests on synthetic data and real data from five Saccharomyces species our algorithm performs significantly better than four other motif-finding algorithms, including algorithms that also take phylogeny into account. Our results also show that, in contrast to the other algorithms, PhyloGibbs can make realistic estimates of the reliability of its predictions. Our tests suggest that, running on the five-species multiple alignment of a single gene's upstream region, PhyloGibbs on average recovers over 50% of all binding sites in S. cerevisiae at a specificity of about 50%, and 33% of all binding sites at a specificity of about 85%. We also tested PhyloGibbs on collections of multiple alignments of intergenic regions that were recently annotated, based on ChIP-on-chip data, to contain binding sites for the same TF. We compared PhyloGibbs's results with the previous analysis of these data using six other motif-finding algorithms. For 16 of 21 TFs for which all other motif-finding methods failed to find a significant motif, PhyloGibbs did recover a motif that matches the literature consensus. In 11 cases where there was disagreement in the results we compiled lists of known target genes from the literature, and found that running PhyloGibbs on their regulatory regions yielded a binding motif matching the literature consensus in all but one of the cases. Interestingly, these literature gene lists had little overlap with the targets annotated based on the ChIP-on-chip data. The PhyloGibbs code can be downloaded from http://www.biozentrum.unibas.ch/~nimwegen/cgi-bin/phylogibbs.cgi or http://www.imsc.res.in/~rsidd/phylogibbs. The full set of predicted sites from our tests on yeast are available at http://www.swissregulon.unibas.ch. Computational discovery of regulatory sites in intergenic DNA is one of the central problems in bioinformatics. Up until recently motif finders would typically take one of the following two general approaches. Given a known set of co-regulated genes, one searches their promoter regions for significantly overrepresented sequence motifs. Alternatively, in a “phylogenetic footprinting” approach one searches multiple alignments of orthologous intergenic regions for short segments that are significantly more conserved than expected based on the phylogeny of the species. In this work the authors present an algorithm, PhyloGibbs, that combines these two approaches into one integrated Bayesian framework. The algorithm searches over all ways in which an arbitrary number of binding sites for an arbitrary number of transcription factors can be assigned to arbitrary collections of multiple sequence alignments while taking into account the phylogenetic relations between the sequences. The authors perform a number of tests on synthetic data and real data from Saccharomyces genomes in which PhyloGibbs significantly outperforms other existing methods. Finally, a novel anneal-and-track strategy allows PhyloGibbs to make accurate estimates of the reliability of its predictions.