ABySS: A parallel assembler for short read sequence data

ABySS: A parallel assembler for short read sequence data
复制标题

DOI:
10.1101/gr.089532.108
复制
发表时间:
2009-06-01
期刊:
影响因子:
7
通讯作者:
Birol, Inanc
Birol, Inanc
中科院分区:
生物学1区
文献类型:
--
作者:
Simpson, Jared T.;Wong, Kim;Birol, Inanc

文献摘要

被引文献

相似文献

大规模平行脱氧核糖核酸(DNA)测序仪器的广泛采用,促使了近年来短读组装算法的发展。现有工具的一个共同缺点是它们无法有效地组装大规模测序项目产生的大量数据,例如对人类个体基因组进行测序以编目自然遗传变异。为了解决这个限制,我们开发了ABySS (Assembly By Short Sequences),一个并行序列汇编器。为了展示我们软件的能力,我们从Illumina公司公开发布的一名非洲男性基因组中收集了35亿对末端读数。共构建了约276万个长度为>= 100碱基对(bp)的contigs, N50大小为1499 bp,占参考人类基因组的68%。对这些序列的分析发现了人类参考序列中不存在的多态性和新序列,并通过与替代的人类序列和其他灵长类基因组比对来验证。
Widespread adoption of massively parallel deoxyribonucleic acid (DNA) sequencing instruments has prompted the recent development of de novo short read assembly algorithms. A common shortcoming of the available tools is their inability to efficiently assemble vast amounts of data generated from large-scale sequencing projects, such as the sequencing of individual human genomes to catalog natural genetic variation. To address this limitation, we developed ABySS (Assembly By Short Sequences), a parallelized sequence assembler. As a demonstration of the capability of our software, we assembled 3.5 billion paired-end reads from the genome of an African male publicly released by Illumina, Inc. Approximately 2.76 million contigs >= 100 base pairs (bp) in length were created with an N50 size of 1499 bp, representing 68% of the reference human genome. Analysis of these contigs identified polymorphic and novel sequences not present in the human reference assembly, which were validated by alignment to alternate human assemblies and to other primate genomes.