PromoterExplorer: an effective promoter identification method based on the AdaBoost algorithm

PromoterExplorer: an effective promoter identification method based on the AdaBoost algorithm
复制标题

DOI:
10.1093/bioinformatics/btl482
复制
发表时间:
2006-11-15
期刊:
影响因子:
5.8
通讯作者:
Yan, Hong
Yan, Hong
中科院分区:
生物学3区
文献类型:
--
作者:
Xie, Xudong;Wu, Shuanhu;Yan, Hong

文献摘要

被引文献

相似文献

动机:启动子预测对基因调控分析具有重要意义。虽然文献中已经报道了许多启动子预测算法,但预测精度的显著提高仍然是一个挑战。本文提出了一种有效的启动子识别算法——PromoterExplorer。在我们的方法中,我们分析了各种特征的不同作用,即五聚体的局部分布、位置CpG岛特征和数字化DNA序列,然后将它们结合起来构建一个高维输入向量。采用基于AdaBoost的级联学习过程,选择最具“信息量”或“判别性”的特征构建一系列弱分类器,将这些弱分类器组合成一个强分类器,从而获得更好的性能。用于识别的级联结构也可以减少误报。结果:PromoterExplorer基于来自EPD、DBTSS、GenBank和人类22号染色体等不同数据库的大量DNA序列进行了测试。实验结果表明,该方法可以获得一致的、令人满意的性能。联系人:h.yan@cityu.edu.hk。
Motivation: Promoter prediction is important for the analysis of gene regulations. Although a number of promoter prediction algorithms have been reported in literature, significant improvement in prediction accuracy remains a challenge. In this paper, an effective promoter identification algorithm, which is called PromoterExplorer, is proposed. In our approach, we analyze the different roles of various features, that is, local distribution of pentamers, positional CpG island features and digitized DNA sequence, and then combine them to build a high-dimensional input vector. A cascade AdaBoost- based learning procedure is adopted to select the most 'informative' or 'discriminating' features to build a sequence of weak classifiers, which are combined to form a strong classifier so as to achieve a better performance. The cascade structure used for identification can also reduce the false positive.Results: PromoterExplorer is tested based on large- scale DNA sequences from different databases, including the EPD, DBTSS, GenBank and human chromosome 22. Experimental results show that consistent and promising performance can be achieved.Contact: h.yan@cityu.edu.hk.