Maximal Words in Sequence Comparisons Based on Subword Composition

Maximal Words in Sequence Comparisons Based on Subword Composition
复制标题

基于子词组合的序列比较中最大词数

DOI:
--
复制
发表时间:
2010
期刊:
Algorithms and Applications
影响因子:
--
通讯作者:
A. Apostolico
A. Apostolico
中科院分区:
--
文献类型:
--
作者:
A. Apostolico

文献摘要

参考文献

被引文献

相似文献

序列相似性和距离的措施或多或少明确地基于子词组成吸引了越来越多的兴趣,密集的应用程序,如大规模的文件分类和全基因组分子分类。这种措施的统一特点是在一些基本概念的相对压缩性,从而两个类似的序列预计将共享更多的共同子串比两个遥远的。本文回顾了一些基于子字组成的序列比较方法,并提出它们的共同点可能最终存在于特殊的子字类中,其性质以有趣的方式与流行的子字树和图的结构产生共鸣。
Measures of sequence similarity and distance based more or less explicitly on subword composition are attracting an increasing interest driven by intensive applications such as massive document classification and genome-wide molecular taxonomy. A uniform character of such measures is in some underlying notion of relative compressibility, whereby two similar sequences are expected to share a larger number of common substrings than two distant ones. This paper reviews some of the approaches to sequence comparison based on subword composition and suggests that their common denominator may ultimately reside in special classes of subwords, the nature of which resonates in interesting ways with the structure of popular subword trees and graphs.
DOI: 10.1016/s0168-9525(00)89076-9
发表时间: 1995-07
期刊: Trends in genetics : TIG
影响因子: --
作者:
Samuel Kariin;C. Burge
通讯作者: Samuel Kariin;C. Burge