Statistical considerations underpinning an alignment-free sequence comparison method

Statistical considerations underpinning an alignment-free sequence comparison method
复制标题

支持免比对序列比较方法的统计考虑

DOI:
10.1016/j.jkss.2010.02.009
复制
发表时间:
2010
期刊:
影响因子:
--
通讯作者:
Susan R. Wilson
Susan R. Wilson
中科院分区:
--
文献类型:
--
作者:
Junmei Jing;C. J. Burden;S. Forêt;Susan R. Wilson

文献摘要

被引文献

相似文献

D2统计量被定义为两个给定序列之间共享的预先指定长度k的词匹配的数量,最多有t个不匹配。该统计量发现其在生物序列的无干扰比较中的应用。它有两个主要的优势,基于测序的方法比较核苷酸和氨基酸序列,如BLAST(基本的本地比对搜索工具)。这些是(i)D2不假设同源片段是连续的,以及(ii)该算法在计算上非常快,在精确匹配的情况下,运行时间与序列的大小成比例。这篇综述文章总结了迄今为止的结果,确定了一系列生物相关参数的分布特性的theD2统计,介绍了现有的应用程序的方法,并概述了未来的研究方向。
TheD2statistic is defined as the number of word matches of prespecified lengthk, with up totmismatches, shared between two given sequences. This statistic finds its application in alignment-free comparisons of biological sequences. It has two main advantages over alignment-based methods for nucleotide and amino-acid sequence comparisons, such as BLAST (basic local alignment search tool). These are (i)D2does not assume that homologous segments are contiguous, and (ii) the algorithm is computationally extremely fast, the runtime being proportional to the size of the sequences in the case of exact matches. This review article summarises results to date on determining the distributional properties of theD2statistic for a range of biologically relevant parameters, describes existing applications of the method, and outlines future research directions.