LASAGNA: a novel algorithm for transcription factor binding site alignment.

LASAGNA: a novel algorithm for transcription factor binding site alignment.
复制标题

DOI:
10.1186/1471-2105-14-108
复制
发表时间:
2013-03-24
期刊:
影响因子:
3
通讯作者:
Huang CH
Huang CH
中科院分区:
生物学4区
文献类型:
--
作者:
Lee C;Huang CH

文献摘要

参考文献

被引文献

相似文献

科学家经常扫描DNA序列中的转录因子(Tf)结合位点(TFBs)。大多数可用的工具依赖于由排列的结合位点构建的位置特定评分矩阵(PSSM)。由于用于获得TFBs的分析的分辨率,TRANSFAC、ORegAnno和Pazar等数据库存储包含TF结合位点的未对齐的可变长度DNA片段。这些DNA片段需要比对才能构建PSSM。虽然TRANSFAC数据库提供了TF的评分矩阵,但公开发布的TF中近78%没有可用的矩阵。由于对TFBS比对算法的研究一直是有限的,因此非常希望有一种针对TFBS的比对算法。我们设计了一种新的算法--千层面,该算法能够感知输入TFBSS的长度,并利用位置相关性。对TRANSFAC数据库中5个物种的189个转录因子的结果表明,我们的方法的性能明显优于ClustaW2和Meme。我们进一步比较了依赖于千层面的PSSM方法和无对齐的TFBS搜索方法。对可以在基因组中定位结合位点的89个TF的结果表明,在固定的召回率下,我们的方法明显更准确。最后,我们描述了千层面,一种更复杂的版本,用于CHIP(染色质免疫沉淀)实验。在每个序列一个的模型下,它在发现芯片序列峰值序列中的基序方面表现出与Meme相当的性能。我们得出结论:千层面算法在比对可变长度结合位点方面是简单有效的。它已经集成到一个用户友好的网络工具中,用于TFBS搜索和可视化,称为千层面搜索。该工具目前分别在TRANSFAC公共数据库(版本7.0)和ORegAnno数据库(08Nov10转储)中存储从TFBS构建的189个TF和133个TF的预计算PSSM模型。该网络工具可在http://biogrid.engr.uconn.edu/lasagna_search/.上获得。
Scientists routinely scan DNA sequences for transcription factor (TF) binding sites (TFBSs). Most of the available tools rely on position-specific scoring matrices (PSSMs) constructed from aligned binding sites. Because of the resolutions of assays used to obtain TFBSs, databases such as TRANSFAC, ORegAnno and PAZAR store unaligned variable-length DNA segments containing binding sites of a TF. These DNA segments need to be aligned to build a PSSM. While the TRANSFAC database provides scoring matrices for TFs, nearly 78% of the TFs in the public release do not have matrices available. As work on TFBS alignment algorithms has been limited, it is highly desirable to have an alignment algorithm tailored to TFBSs. We designed a novel algorithm named LASAGNA, which is aware of the lengths of input TFBSs and utilizes position dependence. Results on 189 TFs of 5 species in the TRANSFAC database showed that our method significantly outperformed ClustalW2 and MEME. We further compared a PSSM method dependent on LASAGNA to an alignment-free TFBS search method. Results on 89 TFs whose binding sites can be located in genomes showed that our method is significantly more precise at fixed recall rates. Finally, we described LASAGNA-ChIP, a more sophisticated version for ChIP (Chromatin immunoprecipitation) experiments. Under the one-per-sequence model, it showed comparable performance with MEME in discovering motifs in ChIP-seq peak sequences. We conclude that the LASAGNA algorithm is simple and effective in aligning variable-length binding sites. It has been integrated into a user-friendly webtool for TFBS search and visualization called LASAGNA-Search. The tool currently stores precomputed PSSM models for 189 TFs and 133 TFs built from TFBSs in the TRANSFAC Public database (release 7.0) and the ORegAnno database (08Nov10 dump), respectively. The webtool is available at http://biogrid.engr.uconn.edu/lasagna_search/.
DOI: 10.1093/nar/gki441
发表时间: 2005-07-01
影响因子: 14.9
作者:
Chekmenev DS;Haid C;Kel AE
通讯作者: Kel AE
DOI: 10.1186/gb-2010-11-2-r19
发表时间: 2010
期刊: Genome biology
影响因子: 12.3
作者:
Georgiev S;Boyle AP;Jayasurya K;Ding X;Mukherjee S;Ohler U
通讯作者: Ohler U
Oreganno:一种开放访问社区驱动的监管注释资源。
DOI: 10.1093/nar/gkm967
发表时间: 2008-01
影响因子: 14.9
作者:
Griffith OL;Montgomery SB;Bernier B;Chu B;Kasaian K;Aerts S;Mahony S;Sleumer MC;Bilenky M;Haeussler M;Griffith M;Gallo SM;Giardine B;Hooghe B;Van Loo P;Blanco E;Ticoll A;Lithwick S;Portales-Casamar E;Donaldson IJ;Robertson G;Wadelius C;De Bleser P;Vlieghe D;Halfon MS;Wasserman W;Hardison R;Bergman CM;Jones SJ;Open Regulatory Annotation Consortium
通讯作者: Open Regulatory Annotation Consortium
DOI: 10.1093/bioinformatics/bti473
发表时间: 2005-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Cartharius, K;Frech, K;Werner, T
通讯作者: Werner, T
DOI: 10.1093/bioinformatics/btm404
发表时间: 2007-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Larkin, M. A.;Blackshields, G.;Higgins, D. G.
通讯作者: Higgins, D. G.