Mining, compressing and classifying with extensible motifs

Mining, compressing and classifying with extensible motifs
复制标题

DOI:
10.1186/1748-7188-1-4
复制
发表时间:
2006-01-01
影响因子:
1
通讯作者:
Parida, Laxmi
Parida, Laxmi
中科院分区:
生物学4区
文献类型:
--
作者:
Apostolico, Alberto;Comin, Matteo;Parida, Laxmi

文献摘要

被引文献

相似文献

背景:最大饱和度的基序模式最初出现在生物分子序列模式发现的背景下,最近证明在数据压缩方案的设计中也是一个有价值的概念。非正式地说,模体是一串断断续续的坚固和狂野的字符,在输入序列或序列家族中或多或少地重复出现。结果:在目前的工作中,考虑了“可扩展的”主题,使得每个空位序列都具有一定的弹性,从而可以将相同的模式拉伸以适应与所有实体字符匹配但长度不同的源片段。然后描述了这一概念的几个应用。在通过文本替换进行数据压缩的应用中,可扩展主题被认为可以节省码本的大小,从而改善压缩。结论:基于可扩展模体的离线压缩可以很好地用于生物序列的压缩和分类。
Background: Motif patterns of maximal saturation emerged originally in contexts of pattern discovery in biomolecular sequences and have recently proven a valuable notion also in the design of data compression schemes. Informally, a motif is a string of intermittently solid and wild characters that recurs more or less frequently in an input sequence or family of sequences. Motif discovery techniques and tools tend to be computationally imposing, however, special classes of "rigid" motifs have been identified of which the discovery is affordable in low polynomial time.Results: In the present work, "extensible" motifs are considered such that each sequence of gaps comes endowed with some elasticity, whereby the same pattern may be stretched to fit segments of the source that match all the solid characters but are otherwise of different lengths. A few applications of this notion are then described. In applications of data compression by textual substitution, extensible motifs are seen to bring savings on the size of the codebook, and hence to improve compression. In germane contexts, in which compressibility is used in its dual role as a basis for structural inference and classification, extensible motifs are seen to support unsupervised classification and phylogeny reconstruction.Conclusion: Off-line compression based on extensible motifs can be used advantageously to compress and classify biological sequences.