Statistical considerations underpinning an alignment-free sequence comparison method
Statistical considerations underpinning an alignment-free sequence comparison method
复制标题
支持免比对序列比较方法的统计考虑
DOI:
10.1016/j.jkss.2010.02.009
复制
发表时间:
2010
期刊:
影响因子:
--
通讯作者:
Susan R. Wilson
中科院分区:
文献类型:
--
作者:
Junmei Jing;C. J. Burden;S. Forêt;Susan R. Wilson
TheD2statistic is defined as the number of word matches of prespecified lengthk, with up totmismatches, shared between two given sequences. This statistic finds its application in alignment-free comparisons of biological sequences. It has two main advantages over alignment-based methods for nucleotide and amino-acid sequence comparisons, such as BLAST (basic local alignment search tool). These are (i)D2does not assume that homologous segments are contiguous, and (ii) the algorithm is computationally extremely fast, the runtime being proportional to the size of the sequences in the case of exact matches. This review article summarises results to date on determining the distributional properties of theD2statistic for a range of biologically relevant parameters, describes existing applications of the method, and outlines future research directions.