A novel ensemble learning method for de novo computational identification of DNA binding sites.

A novel ensemble learning method for de novo computational identification of DNA binding sites.
复制标题

DOI:
10.1186/1471-2105-8-249
复制
发表时间:
2007-07-12
期刊:
影响因子:
3
通讯作者:
Gross RH
Gross RH
中科院分区:
生物学4区
文献类型:
--
作者:
Chakravarty A;Carlson JM;Khetani RS;Gross RH

文献摘要

参考文献

被引文献

相似文献

尽管基序表示和搜索算法具有多样性,但转录因子结合位点的从头计算识别仍然受到现有算法的有限准确性以及需要用户指定的描述所寻找基序的输入参数的限制。我们提出了一种新颖的集成学习方法 SCOPE,该方法基于这样的假设:转录因子结合位点属于三大类基序之一:非简并基序、简并基序和缺口基序。 SCOPE 采用统一的评分指标来结合三种主题查找算法的结果,每个算法旨在发现这些类别的主题之一。我们发现 SCOPE 对来自 4 个物种的 78 个经过实验表征的调节子的性能比其组件算法有了实质性的、统计上显着的改进。在相同的数据集上,SCOPE 的表现优于多种现有的主题发现算法,具有统计学上的显着优势。 SCOPE 证明,结合多个集中的主题发现算法可以显着提高性能。通过构建无需用户定义参数即可有效搜索基序的组件,SCOPE 只需要一组上游序列和物种名称作为输入,使其成为非专家用户的实用选择。用户友好的 Web 界面、Java 源代码和可执行文件可在 处获得。
Despite the diversity of motif representations and search algorithms, the de novo computational identification of transcription factor binding sites remains constrained by the limited accuracy of existing algorithms and the need for user-specified input parameters that describe the motif being sought. We present a novel ensemble learning method, SCOPE, that is based on the assumption that transcription factor binding sites belong to one of three broad classes of motifs: non-degenerate, degenerate and gapped motifs. SCOPE employs a unified scoring metric to combine the results from three motif finding algorithms each aimed at the discovery of one of these classes of motifs. We found that SCOPE's performance on 78 experimentally characterized regulons from four species was a substantial and statistically significant improvement over that of its component algorithms. SCOPE outperformed a broad range of existing motif discovery algorithms on the same dataset by a statistically significant margin. SCOPE demonstrates that combining multiple, focused motif discovery algorithms can provide a significant gain in performance. By building on components that efficiently search for motifs without user-defined parameters, SCOPE requires as input only a set of upstream sequences and a species designation, making it a practical choice for non-expert users. A user-friendly web interface, Java source code and executables are available at .
DOI: 10.1186/gb-2004-5-9-r61
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Berman BP;Pfeiffer BD;Laverty TR;Salzberg SL;Rubin GM;Eisen MB;Celniker SE
通讯作者: Celniker SE
DOI: 10.1073/pnas.231608898
发表时间: 2002-01-22
影响因子: 11.1
作者:
Berman, BP;Nibu, Y;Eisen, MB
通讯作者: Eisen, MB
DOI: 10.1186/1471-2105-8-249
发表时间: 2007-07-12
期刊: BMC bioinformatics
影响因子: 3
作者:
Chakravarty A;Carlson JM;Khetani RS;Gross RH
通讯作者: Gross RH
DOI: 10.1186/1471-2105-7-254
发表时间: 2006-05-15
期刊: BMC bioinformatics
影响因子: 3
作者:
Carlson JM;Chakravarty A;Khetani RS;Gross RH
通讯作者: Gross RH
DOI: 10.1093/nar/gkl372
发表时间: 2006
影响因子: 14.9
作者:
GuhaThakurta D
通讯作者: GuhaThakurta D