Combinatorial pattern discovery in biological sequences: the TEIRESIAS algorithm

Combinatorial pattern discovery in biological sequences: the TEIRESIAS algorithm
复制标题

DOI:
10.1093/bioinformatics/14.1.55
复制
发表时间:
1998-01-01
期刊:
影响因子:
5.8
通讯作者:
Floratos, A
Floratos, A
中科院分区:
生物学3区
文献类型:
--
作者:
Rigoutsos, I;Floratos, A

文献摘要

被引文献

相似文献

动机:在生物序列中发现基序是一个重要问题。 结果:本文提出了一种在生物序列中发现刚性模式(基序)的新算法。我们的方法本质上是组合性的,能够产生至少在(用户定义的)最小数量的序列中出现的所有模式,而且通过避免对整个模式空间进行枚举,它能够非常高效。此外,所报告的模式是最大的:任何所报告的模式都不能变得更具体,同时仍在输入序列内的完全相同位置出现。所提出的方法的有效性在一些测试案例中得到了展示,这些测试案例旨在:(i)通过发现先前报告的模式来验证该方法;(ii)展示针对所考虑的序列自动识别高度选择性模式的能力。最后,实验分析表明该算法对输出敏感,即其运行时间与所生成的输出大小近似呈线性关系。
Motivation: The discovery of motifs in biological sequences is an important problem.Results: This paper presents a new algorithm for the discovery of rigid patterns (motifs) in biological sequences. Our method is combinatorial in nature and able to produce all patterns that appear in at least a (user-defined) minimum number of sequences, yet it manages to be very efficient by avoiding the enumeration of the entire pattern space. Furthermore, the reported patterns are maximal: any reported pattern cannot be made more specific and still keep on appearing at the exact same positions within the input sequences. The effectiveness of the proposed approach is showcased on a number of test cases which aim to: (i) validate the approach through the discovery of previously reported patterns; (ii) demonstrate the capability to identify automatically highly selective patterns particular to the sequences under consideration. Finally, experimental analysis indicates that the algorithm is output sensitive, i.e. its running time is quasi-linear to the size of the generated output.