Discovering sequence motifs with arbitrary insertions and deletions.

Discovering sequence motifs with arbitrary insertions and deletions.
复制标题

发现具有任意插入和删除的序列基序。

DOI:
10.1371/journal.pcbi.1000071
复制
发表时间:
2008-05-09
影响因子:
4.3
通讯作者:
Bailey, Timothy L.
Bailey, Timothy L.
中科院分区:
生物学2区
文献类型:
--
作者:
Frith, Martin C.;Saunders, Neil F. W.;Kobe, Bostjan;Bailey, Timothy L.

文献摘要

参考文献

被引文献

相似文献

生物学信息编码于分子序列中:解读这种编码仍然是一项重大的科学挑战。DNA、RNA和蛋白质序列的功能区域往往呈现出具有特征但很微妙的基序;因此,通过计算在序列中发现基序是一个基础性且被广泛研究的问题。然而,目前大多数算法不允许基序内部存在插入或缺失(插入缺失),而少数允许的算法又存在其他局限性。我们提出了一种方法,GLAM2(基序的含间隙局部比对),用于以一种完全通用的方式发现允许插入缺失的基序,以及一种配套方法GLAM2SCAN,用于使用此类基序搜索序列数据库。GLAM2是无间隙吉布斯采样算法的一种推广。它从PROSITE数据库中重新发现可变宽度的蛋白质基序,其准确性明显高于替代方法PRATT和SAM - T2K。此外,它对ELM数据库中的蛋白质基序进行了有效的优化:在某些情况下,优化后的基序比原始的ELM正则表达式产生的过度预测要少几个数量级。GLAM2在BAliBASE多序列比对基准测试中表现良好,并且对于具有N端和C端延伸的“基序样”比对,可能优于领先的多序列比对方法。最后,我们展示了使用GLAM2发现蛋白激酶底物基序以及仅含LIM的转录调控复合物的含间隙DNA基序:使用GLAM2SCAN,我们确定了后者有希望的靶点。GLAM2对于短蛋白质基序尤其有前景,它应该会提高我们识别蛋白质切割位点、相互作用位点、翻译后修饰附着位点等的能力,这些位点是生物学的重要基础。它对于DNA和RNA中任意含间隙的基序可能同样有用,尽管目前此类基序的实例较少。GLAM2是公共领域软件,可从http://bioinformatics.org.au/glam2下载。 近几十年来,科学家们从众多生物体中提取了基因序列——DNA、RNA和蛋白质序列。这些序列包含了这些生物体构建和运作的信息,但到目前为止我们大多还无法解读它们。人们早就知道这些序列包含许多种“基序”,即与特定生物学功能相关的重复模式。因此,大量研究致力于开发计算机算法,用于自动发现序列中微妙的、重复出现的基序。然而,先前的算法搜索的是固定的基序,其实例仅通过替换而变化,而不是通过插入或缺失。实际的基序是灵活的,并且确实会因插入和缺失而变化。这项研究描述了一种用于发现基序的新型计算机算法,该算法允许任意的插入和缺失。这种算法能够发现真实的、灵活的基序,并且应该能够帮助我们确定许多生物分子的功能。
Biology is encoded in molecular sequences: deciphering this encoding remains a grand scientific challenge. Functional regions of DNA, RNA, and protein sequences often exhibit characteristic but subtle motifs; thus, computational discovery of motifs in sequences is a fundamental and much-studied problem. However, most current algorithms do not allow for insertions or deletions (indels) within motifs, and the few that do have other limitations. We present a method, GLAM2 (Gapped Local Alignment of Motifs), for discovering motifs allowing indels in a fully general manner, and a companion method GLAM2SCAN for searching sequence databases using such motifs. glam2 is a generalization of the gapless Gibbs sampling algorithm. It re-discovers variable-width protein motifs from the PROSITE database significantly more accurately than the alternative methods PRATT and SAM-T2K. Furthermore, it usefully refines protein motifs from the ELM database: in some cases, the refined motifs make orders of magnitude fewer overpredictions than the original ELM regular expressions. GLAM2 performs respectably on the BAliBASE multiple alignment benchmark, and may be superior to leading multiple alignment methods for “motif-like” alignments with N- and C-terminal extensions. Finally, we demonstrate the use of GLAM2 to discover protein kinase substrate motifs and a gapped DNA motif for the LIM-only transcriptional regulatory complex: using GLAM2SCAN, we identify promising targets for the latter. GLAM2 is especially promising for short protein motifs, and it should improve our ability to identify the protein cleavage sites, interaction sites, post-translational modification attachment sites, etc., that underlie much of biology. It may be equally useful for arbitrarily gapped motifs in DNA and RNA, although fewer examples of such motifs are known at present. GLAM2 is public domain software, available for download at http://bioinformatics.org.au/glam2. In recent decades, scientists have extracted genetic sequences—DNA, RNA, and protein sequences—from numerous organisms. These sequences hold the information for the construction and functioning of these organisms, but as yet we are mostly unable to read them. It has long been known that these sequences contain many kinds of “motifs”, i.e. re-occurring patterns, associated with specific biological functions. Thus, much research has been devoted to computer algorithms for automatically discovering subtle, recurring motifs in sequences. However, previous algorithms search for rigid motifs whose instances vary only by substitutions, and not by insertions or deletions. Real motifs are flexible, and do vary by insertions and deletions. This study describes a new computer algorithm for discovering motifs, which allows for arbitrary insertions and deletions. This algorithm can discover real, flexible motifs, and should be able to help us determine the functions of many biological molecules.
DOI: 10.1093/nar/gkh169
发表时间: 2004-01-01
影响因子: 14.9
作者:
Frith, MC;Hansen, U;Weng, ZP
通讯作者: Weng, ZP
DOI: 10.1016/j.bbapap.2005.07.036
发表时间: 2005-12-30
影响因子: 3.2
作者:
Kobe, B;Kampmann, T;Worth, RI
通讯作者: Worth, RI
DOI: 10.1186/1471-2105-8-381
发表时间: 2007-10-11
期刊: BMC bioinformatics
影响因子: 3
作者:
Caffrey DR;Dana PH;Mathur V;Ocano M;Hong EJ;Wang YE;Somaroo S;Caffrey BE;Potluri S;Huang ES
通讯作者: Huang ES
DOI: 10.1128/mcb.24.4.1439-1452.2004
发表时间: 2004-02-01
影响因子: 5.3
作者:
Lahlil, R;Lécuyer, E;Hoang, T
通讯作者: Hoang, T
DOI: 10.1093/bioinformatics/bth088
发表时间: 2004-06-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Beissbarth, T;Speed, TP
通讯作者: Speed, TP