Bounded search for de novo identification of degenerate cis-regulatory elements.

Bounded search for de novo identification of degenerate cis-regulatory elements.
复制标题

DOI:
10.1186/1471-2105-7-254
复制
发表时间:
2006-05-15
期刊:
影响因子:
3
通讯作者:
Gross RH
Gross RH
中科院分区:
生物学4区
文献类型:
--
作者:
Carlson JM;Chakravarty A;Khetani RS;Gross RH

文献摘要

参考文献

被引文献

相似文献

从理论上讲,共调节基因上游区域中统计学上过度表达的序列的鉴定应该允许鉴定潜在的顺式调节元件。然而,在实践中,许多顺式调控元件是高度简并的,排除了使用详尽的字数统计策略来识别它们。虽然有许多方法可以使用位置权重矩阵来推断基础分布,但最近的研究表明,模型中固有的独立性假设以及无法达到全局最优值限制了这种方法。在本文中,我们报告PRISM,一个退化的基序发现者,利用一组结合位点的统计意义和个人的结合位点之间的关系。PRISM首先识别过度代表的非简并共有基序,然后迭代地将每个基序松弛成高分简并基序。这种方法不需要可调参数,从而有助于进行无偏的性能比较。因此,我们比较PRISM的性能与9个流行的主题发现28个特征良好的S。酿酒酵母调节子PRISM始终优于所有其他程序。最后,我们使用PRISM来预测未表征的调节子的结合位点。我们的研究结果支持酵母细胞周期转录因子Stb1,其结合位点尚未通过实验确定的拟议的作用机制。结合位点的统计测量值和作为一个整体的集合之间的关系导致识别蛋白质结合的各种顺式调节元件的简单方法。这种方法利用了字数统计的优点,因为位置依赖性被隐式地考虑,并且更容易避免局部最优。虽然我们牺牲保证的最优性,以防止指数爆破的穷举搜索,我们证明了错误是有界的,实验表明,性能上级其他方法。该算法的Java实现可以从我们的Web服务器下载。
The identification of statistically overrepresented sequences in the upstream regions of coregulated genes should theoretically permit the identification of potential cis-regulatory elements. However, in practice many cis-regulatory elements are highly degenerate, precluding the use of an exhaustive word-counting strategy for their identification. While numerous methods exist for inferring base distributions using a position weight matrix, recent studies suggest that the independence assumptions inherent in the model, as well as the inability to reach a global optimum, limit this approach. In this paper, we report PRISM, a degenerate motif finder that leverages the relationship between the statistical significance of a set of binding sites and that of the individual binding sites. PRISM first identifies overrepresented, non-degenerate consensus motifs, then iteratively relaxes each one into a high-scoring degenerate motif. This approach requires no tunable parameters, thereby lending itself to unbiased performance comparisons. We therefore compare PRISM's performance against nine popular motif finders on 28 well-characterized S. cerevisiae regulons. PRISM consistently outperforms all other programs. Finally, we use PRISM to predict the binding sites of uncharacterized regulons. Our results support a proposed mechanism of action for the yeast cell-cycle transcription factor Stb1, whose binding site has not been determined experimentally. The relationship between statistical measures of the binding sites and the set as a whole leads to a simple means of identifying the diverse range of cis-regulatory elements to which a protein binds. This approach leverages the advantages of word-counting, in that position dependencies are implicitly accounted for and local optima are more easily avoided. While we sacrifice guaranteed optimality to prevent the exponential blowup of exhaustive search, we prove that the error is bounded and experimentally show that the performance is superior to other methods. A Java implementation of this algorithm can be downloaded from our web server at .
DOI: 10.1007/bf00993379
发表时间: 1995-10-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
BAILEY, TL;ELKAN, C
通讯作者: ELKAN, C
DOI: 10.1073/pnas.94.11.5617
发表时间: 1997-05-27
影响因子: 11.1
作者:
Isalan, M;Choo, Y;Klug, A
通讯作者: Klug, A
DOI: 10.1089/10665270252935430
发表时间: 2002-01-01
影响因子: 1.7
作者:
Buhler, J;Tompa, M
通讯作者: Tompa, M
DOI: 10.1002/bies.10073
发表时间: 2002-05-01
期刊: BIOESSAYS
影响因子: 4
作者:
Benos, PV;Lapedes, AS;Stormo, GD
通讯作者: Stormo, GD
DOI: 10.1016/s0092-8674(01)00494-9
发表时间: 2001-09-21
期刊: CELL
影响因子: 64.5
作者:
Simon, I;Barnett, J;Young, RA
通讯作者: Young, RA