An Efficient Algorithm for Discovering Motifs in Large DNA Data Sets

An Efficient Algorithm for Discovering Motifs in Large DNA Data Sets
复制标题

一种在大型 DNA 数据集中发现基序的有效算法

DOI:
10.1109/tnb.2015.2421340
复制
发表时间:
2015-07-01
影响因子:
3.9
通讯作者:
Huan, Jun
Huan, Jun
中科院分区:
生物学3区
文献类型:
--
作者:
Yu, Qiang;Huo, Hongwei;Huan, Jun

文献摘要

被引文献

相似文献

在过去的十年里,种植(L,d)基序的发现已经成功地用于在数十个启动子序列中定位转录因子结合位点。然而,在包含数以千计的输入序列的下一代测序(CHIP-SEQ)数据集中,(L,d)基序的识别工作还做得不够,从而给在合理的时间内进行良好的识别带来了新的挑战。为了满足这一需求,我们提出了一种新的种植(L,d)基序发现算法MCES,该算法通过挖掘和组合出现的子串来识别基序。特别是,为了处理更大的数据集,我们设计了一种基于MapReduce的策略来分布式挖掘新出现的子串。在模拟数据上的实验结果表明:i)MCES能够在数千到数百万个输入序列中高效地识别(L,d)基序,并且运行速度快于最新的(L,d)基序发现算法,如F-Motif和TraverStringsR;ii)MCES能够识别长度未知的基序,并且比竞争算法Cisfinder具有更好的识别精度。同时,在真实数据集上测试了MCES的有效性。MCES可在http://sites.google.com/site/feqond/mces.上免费获得
The planted (l,d) motif discovery has been successfully used to locate transcription factor binding sites in dozens of promoter sequences over the past decade. However, there has not been enough work done in identifying (l,d) motifs in the next-generation sequencing (ChIP-seq) data sets, which contain thousands of input sequences and thereby bring new challenge to make a good identification in reasonable time. To cater this need, we propose a new planted (l,d) motif discovery algorithm named MCES, which identifies motifs by mining and combining emerging substrings. Specially, to handle larger data sets, we design a MapReduce-based strategy to mine emerging substrings distributedly. Experimental results on the simulated data show that i) MCES is able to identify (l,d) motifs efficiently and effectively in thousands to millions of input sequences, and runs faster than the state-of-the-art (l,d) motif discovery algorithms, such as F-motif and TraverStringsR; ii) MCES is able to identify motifs without known lengths, and has a better identification accuracy than the competing algorithm CisFinder. Also, the validity of MCES is tested on real data sets. MCES is freely available at http://sites.google.com/site/feqond/mces.