Assembling genomes on large-scale parallel computers

Assembling genomes on large-scale parallel computers
复制标题

DOI:
10.1016/j.jpdc.2007.05.014
复制
发表时间:
2007-12-01
影响因子:
3.8
通讯作者:
Aluru, S.
Aluru, S.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Kalyanaraman, A.;Emrich, S. J.;Aluru, S.

文献摘要

被引文献

相似文献

从数千万个短基因组片段组装大基因组在计算上要求很高,需要数百千兆字节的存储器和数万个CPU小时。高通量测序技术、新的基因富集测序策略和环境样品的集体测序的出现进一步加剧了这种情况。在本文中,我们提出了第一个大规模并行基因组组装框架。我们的方法的独特功能包括空间效率和按需算法,只消耗线性空间,和战略,以减少昂贵的成对序列比对的数量,同时保持装配质量。作为玉米基因组测序中正在进行的努力的一部分,我们将我们的组装框架应用于包含基因富集和随机鸟枪序列的混合物的基因组数据。我们报告称,在IBM BlueGene/L超级计算机的1024个处理器上,在2小时内将超过160万个片段4(总大小超过12.5亿个核苷酸)划分为基因组岛。我们还证明了传统的全基因组鸟枪测序和组装环境序列的方法的有效性。(c)2007年爱思唯尔公司所有的战斗保留。
Assembly of large genomes from tens of millions of short genomic fragments is computationally demanding requiring hundreds of gigabytes of memory and tens of thousands of CPU hours. The advent of high throughput sequencing technologies, new gene-enrichment sequencing strategies, and collective sequencing of environmental samples further exacerbate this situation. In this paper, we present the first massively parallel genome assembly framework. The unique features of our approach include space-efficient and on-demand algorithms that consume only linear space, and strategies to reduce the number of expensive pairwise sequence alignments while maintaining assembly quality. Developed as pan of the ongoing efforts in maize genome sequencing, we applied our assembly framework to genomic data containing a mixture of gene enriched and random shotgun sequences. We report the partitioning of more than 1.6 million fragments 4 over 1.25 billion nucleotides total size into genomic islands in under 2 h on 1024 processors of an IBM BlueGene/L supercomputer. We also demonstrate the effectiveness of the proposed approach for traditional whole genome shotgun sequencing and assembly of environmental sequences. (c) 2007 Elsevier Inc. All fights reserved.