TFBS identification based on genetic algorithm with combined representations and adaptive post-processing

TFBS identification based on genetic algorithm with combined representations and adaptive post-processing
复制标题

DOI:
10.1093/bioinformatics/btm606
复制
发表时间:
2008-02-01
期刊:
影响因子:
5.8
通讯作者:
Lee, Kin-Hong
Lee, Kin-Hong
中科院分区:
生物学3区
文献类型:
--
作者:
Chan, Tak-Ming;Leung, Kwong-Sak;Lee, Kin-Hong

文献摘要

被引文献

相似文献

动机:转录因子结合位点(TFBS)的识别在破译基因调控机制中发挥着重要作用。最近,GAME(一种基于遗传算法(GA)的迭代后处理方法)在 TFBS 识别中表现出了卓越的性能。然而GAME中的基本遗传算法设计并不精细,在实际问题中可能会陷入局部最优。特征算子仅应用于后处理,但最终性能很大程度上取决于遗传算法的输出。因此,通过在遗传算法中引入更先进的表示和新颖的算子,以及以自适应方式设计后处理,可以提高整个算法的有效性和效率。结果:我们提出了一种新颖的框架GALF-P,由局部过滤遗传算法(GALF)和自适应后处理技术(-P)组成,以实现TFBS识别的有效性和效率。 GALF 结合了当前 GA 中分别使用的位置主导和共识主导的表示形式,并采用一种新颖的局部过滤算子,在 GA 的进化过程中有效地消除个体中的误报。预选择用于保持多样性并避免局部最优。开发了具有自适应添加和删除功能的后处理,以处理每个序列具有任意数量实例的一般情况。 GALF-P 在具有困难场景的合成数据集和真实测试数据集上显示出优于 GAME、MEME、BioProspector 和 BioOptimizer 的性能。进一步与当前最先进的方法 GAME 相比,GALF-P 也更加稳健和可靠。
Motivation: Identification of transcription factor binding sites (TFBSs) plays an important role in deciphering the mechanisms of gene regulation. Recently, GAME, a Genetic Algorithm (GA)-based approach with iterative post-processing, has shown superior performance in TFBS identification. However, the basic GA in GAME is not elaborately designed, and may be trapped in local optima in real problems. The feature operators are only applied in the post-processing, but the final performance heavily depends on the GA output. Hence, both effectiveness and efficiency of the overall algorithm can be improved by introducing more advanced representations and novel operators in the GA, as well as designing the post-processing in an adaptive way.Results: We propose a novel framework GALF-P, consisting of Genetic Algorithm with Local Filtering (GALF) and adaptive post-processing techniques (-P), to achieve both effectiveness and efficiency for TFBS identification. GALF combines the position-led and consensus-led representations used separately in current GAs and employs a novel local filtering operator to get rid of false positives within an individual efficiently during the evolutionary process in the GA. Pre-selection is used to maintain diversity and avoid local optima. Post-processing with adaptive adding and removing is developed to handle general cases with arbitrary numbers of instances per sequence. GALF-P shows superior performance to GAME, MEME, BioProspector and BioOptimizer on synthetic datasets with difficult scenarios and real test datasets. GALF-P is also more robust and reliable when further compared with GAME, the current state-of-the-art approach.