A fast weak motif-finding algorithm based on community detection in graphs.

A fast weak motif-finding algorithm based on community detection in graphs.
复制标题

一种基于图中社区检测的快速弱主题发现算法

DOI:
10.1186/1471-2105-14-227
复制
发表时间:
2013-07-17
期刊:
影响因子:
3
通讯作者:
Yu J
Yu J
中科院分区:
生物学4区
文献类型:
--
作者:
Jia C;Carson MB;Yu J

文献摘要

参考文献

被引文献

相似文献

在DNA序列中识别转录因子结合位点(也称为“基序发现”)是理解基因调控的基本步骤。尽管已经开发了许多成功的程序,但由于基因表达/调控的多样性以及结合位点的低特异性,这个问题远未得到解决。最先进的算法都有自身的局限性(例如,寻找长基序时时间或空间复杂度高,识别弱基序精度低,或者存在OOPS限制:每个序列中基序实例仅出现一次),这些限制了它们的应用范围。 在本文中,我们提出了一种新颖且快速的算法,我们称之为TFBSGroup。它基于从图中进行社区检测,用于在ZOMOPS限制(每个序列中基序实例出现零次、一次或多次)下发现长且弱的(l,d)基序,其中l是基序的长度,d是基序实例与基序本身之间的最大突变数。首先,TFBSGroup将序列中的(l,d)基序搜索转换为侧重于图内密集子图的发现。它使用一种快速的社区检测方法来识别这些子图,以获得粗粒度的候选基序。接下来,它在各自的社区内朝着真实基序贪婪地细化这些候选基序。对合成的(l,d)样本的实证研究表明,TFBSGroup非常高效(例如,它能在30秒内找到真实的(18,6)、(24,8)基序)。更重要的是,该算法成功地在从大肠杆菌数据库RegulonDB生成的大量原核启动子数据集中快速识别了基序。该算法还准确地识别了涉及胚胎干细胞多能性和自我更新的12种小鼠转录因子的ChIP - seq数据集中的基序。 我们新颖的启发式算法TFBSGroup能够在ZOMOPS限制下快速识别DNA序列中长且弱的(l,d)基序的近乎精确匹配。它还能够在实际应用中发现基序。TFBSGroup的源代码可从http://bioinformatics.bioengr.uic.edu/TFBSGroup/获取。
Identification of transcription factor binding sites (also called ‘motif discovery’) in DNA sequences is a basic step in understanding genetic regulation. Although many successful programs have been developed, the problem is far from being solved on account of diversity in gene expression/regulation and the low specificity of binding sites. State-of-the-art algorithms have their own constraints (e.g., high time or space complexity for finding long motifs, low precision in identification of weak motifs, or the OOPS constraint: one occurrence of the motif instance per sequence) which limit their scope of application. In this paper, we present a novel and fast algorithm we call TFBSGroup. It is based on community detection from a graph and is used to discover long and weak (l,d) motifs under the ZOMOPS constraint (zero, one or multiple occurrence(s) of the motif instance(s) per sequence), where l is the length of a motif and d is the maximum number of mutations between a motif instance and the motif itself. Firstly, TFBSGroup transforms the (l, d) motif search in sequences to focus on the discovery of dense subgraphs within a graph. It identifies these subgraphs using a fast community detection method for obtaining coarse-grained candidate motifs. Next, it greedily refines these candidate motifs towards the true motif within their own communities. Empirical studies on synthetic (l, d) samples have shown that TFBSGroup is very efficient (e.g., it can find true (18, 6), (24, 8) motifs within 30 seconds). More importantly, the algorithm has succeeded in rapidly identifying motifs in a large data set of prokaryotic promoters generated from the Escherichia coli database RegulonDB. The algorithm has also accurately identified motifs in ChIP-seq data sets for 12 mouse transcription factors involved in ES cell pluripotency and self-renewal. Our novel heuristic algorithm, TFBSGroup, is able to quickly identify nearly exact matches for long and weak (l, d) motifs in DNA sequences under the ZOMOPS constraint. It is also capable of finding motifs in real applications. The source code for TFBSGroup can be obtained from http://bioinformatics.bioengr.uic.edu/TFBSGroup/.
DOI: 10.1089/10665270252935430
发表时间: 2002-01-01
影响因子: 1.7
作者:
Buhler, J;Tompa, M
通讯作者: Tompa, M
DOI: 10.1186/gb-2010-11-2-r19
发表时间: 2010
期刊: Genome biology
影响因子: 12.3
作者:
Georgiev S;Boyle AP;Jayasurya K;Ding X;Mukherjee S;Ohler U
通讯作者: Ohler U
DOI: 10.1093/bioinformatics/bti336
发表时间: 2005-05-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Favorov, AV;Gelfand, MS;Makeev, VJ
通讯作者: Makeev, VJ
DOI: 10.1088/1742-5468/2008/10/p10008
发表时间: 2008-10-01
影响因子: 2.4
作者:
Blondel, Vincent D.;Guillaume, Jean-Loup;Lefebvre, Etienne
通讯作者: Lefebvre, Etienne
DOI: 10.1128/jb.177.17.4872-4880.1995
发表时间: 1995-09-01
影响因子: 3.2
作者:
CUI, YH;WANG, Q;CALVO, JM
通讯作者: CALVO, JM