Adaptive seeds tame genomic sequence comparison

Adaptive seeds tame genomic sequence comparison
复制标题

DOI:
10.1101/gr.113985.110
复制
发表时间:
2011-03-01
期刊:
影响因子:
7
通讯作者:
Frith, Martin C.
Frith, Martin C.
中科院分区:
生物学1区
文献类型:
--
作者:
Kielbasa, Szymon M.;Wan, Raymond;Frith, Martin C.

文献摘要

被引文献

相似文献

分析生物序列的主要方法是将它们相互比较和比对。然而,要比较现代数十亿碱基的DNA数据集仍然很困难。困难是由这些序列的不均匀(寡)核苷酸组成引起的,而不是它们本身的大小。为了解决这个问题,我们修改了标准的种子和扩展方法(e。例如,在一个实施例中,BLAST)来使用自适应种子。自适应种子是根据其稀有性选择的匹配,而不是使用固定长度的匹配。这个方法保证了匹配的数量以及运行时间随序列长度线性增加,而不是二次增加。LAST是我们的自适应种子的开源实现,可以快速和灵敏地比较具有任意非均匀组成的大序列。
The main way of analyzing biological sequences is by comparing and aligning them to each other. It remains difficult, however, to compare modern multi-billionbase DNA data sets. The difficulty is caused by the nonuniform (oligo) nucleotide composition of these sequences, rather than their size per se. To solve this problem, we modified the standard seed-and-extend approach (e. g., BLAST) to use adaptive seeds. Adaptive seeds are matches that are chosen based on their rareness, instead of using fixed-length matches. This method guarantees that the number of matches, and thus the running time, increases linearly, instead of quadratically, with sequence length. LAST, our open source implementation of adaptive seeds, enables fast and sensitive comparison of large sequences with arbitrarily nonuniform composition.