Assembling millions of short DNA sequences using SSAKE

Assembling millions of short DNA sequences using SSAKE
复制标题

DOI:
10.1093/bioinformatics/btl629
复制
发表时间:
2007-02-15
期刊:
影响因子:
5.8
通讯作者:
Holt, Robert A.
Holt, Robert A.
中科院分区:
生物学3区
文献类型:
--
作者:
Warren, Rene L.;Sutton, Granger G.;Holt, Robert A.

文献摘要

被引文献

相似文献

新的DNA测序技术与潜力高达三个数量级的序列吞吐量比传统的桑格测序正在出现。该仪器现在可从Solexa有限公司获得,产生数百万个短DNA序列,每个序列25 nt。由于大基因组中普遍存在的重复序列,以及短序列无法独特和明确地表征它们,短的读取长度限制了从头测序的适用性。然而,考虑到该仪器的测序深度和通量,可以实现高度相同序列的严格组装。我们描述SSAKE,一个工具,积极组装数以百万计的短核苷酸序列,逐步搜索通过前缀树最长可能重叠的任何两个序列之间。SSAKE旨在通过严格地将它们组装成可用于表征新测序目标的连续序列来帮助利用短序列读取的信息。
Novel DNA sequencing technologies with the potential for up to three orders magnitude more sequence throughput than conventional Sanger sequencing are emerging. The instrument now available from Solexa Ltd, produces millions of short DNA sequences of 25 nt each. Due to ubiquitous repeats in large genomes and the inability of short sequences to uniquely and unambiguously characterize them, the short read length limits applicability for de novo sequencing. However, given the sequencing depth and the throughput of this instrument, stringent assembly of highly identical sequences can be achieved. We describe SSAKE, a tool for aggressively assembling millions of short nucleotide sequences by progressively searching through a prefix tree for the longest possible overlap between any two sequences. SSAKE is designed to help leverage the information from short sequence reads by stringently assembling them into contiguous sequences that can be used to characterize novel sequencing targets.