CloudBurst: highly sensitive read mapping with MapReduce.

CloudBurst: highly sensitive read mapping with MapReduce.
复制标题

DOI:
10.1093/bioinformatics/btp236
复制
发表时间:
2009-06-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Schatz MC
Schatz MC
中科院分区:
其他
文献类型:
--
作者:
Schatz MC

文献摘要

参考文献

被引文献

相似文献

动机:新一代DNA测序仪正在产生大量的序列数据,对传统的单处理器读段比对算法提出了前所未有的要求。CloudBurst是一种新的并行读段比对算法,针对将新一代序列数据比对到人类基因组和其他参考基因组进行了优化,可用于包括单核苷酸多态性(SNP)发现、基因分型和个人基因组学在内的多种生物学分析。它以短读段比对程序RMAP为模型,能够报告每个读段的所有比对结果,或者在存在任意数量错配或差异的情况下报告明确的最佳比对结果。这种敏感度可能会耗费极长的时间,但CloudBurst使用开源的Hadoop版MapReduce,利用多个计算节点并行执行。 结果:CloudBurst的运行时间与所比对的读段数量呈线性关系,并且随着处理器数量的增加,加速比接近线性。在24核处理器配置下,CloudBurst比在单核上执行的RMAP快达30倍,同时计算出相同的比对结果集。使用具有96核的更大远程计算云,CloudBurst将性能提高了100多倍,对于涉及将数百万条短读段比对到人类基因组的典型任务,将运行时间从数小时缩短到仅仅几分钟。 可用性:CloudBurst作为使用MapReduce并行化算法的一个模型,可在http://cloudburst - bio.sourceforge.net/获取开源版本。 联系方式:mschatz@umiacs.umd.edu
Motivation: Next-generation DNA sequencing machines are generating an enormous amount of sequence data, placing unprecedented demands on traditional single-processor read-mapping algorithms. CloudBurst is a new parallel read-mapping algorithm optimized for mapping next-generation sequence data to the human genome and other reference genomes, for use in a variety of biological analyses including SNP discovery, genotyping and personal genomics. It is modeled after the short read-mapping program RMAP, and reports either all alignments or the unambiguous best alignment for each read with any number of mismatches or differences. This level of sensitivity could be prohibitively time consuming, but CloudBurst uses the open-source Hadoop implementation of MapReduce to parallelize execution using multiple compute nodes. Results: CloudBurst's running time scales linearly with the number of reads mapped, and with near linear speedup as the number of processors increases. In a 24-processor core configuration, CloudBurst is up to 30 times faster than RMAP executing on a single core, while computing an identical set of alignments. Using a larger remote compute cloud with 96 cores, CloudBurst improved performance by >100-fold, reducing the running time from hours to mere minutes for typical jobs involving mapping of millions of short reads to the human genome. Availability: CloudBurst is available open-source as a model for parallelizing algorithms with MapReduce at http://cloudburst-bio.sourceforge.net/. Contact: mschatz@umiacs.umd.edu
DOI: 10.1186/1471-2105-9-128
发表时间: 2008-02-28
期刊: BMC bioinformatics
影响因子: 3
作者:
Smith AD;Xuan Z;Zhang MQ
通讯作者: Zhang MQ
DOI: 10.1101/gr.078212.108
发表时间: 2008-11-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Li, Heng;Ruan, Jue;Durbin, Richard
通讯作者: Durbin, Richard
亚洲个体的二倍体基因组序列
DOI: 10.1038/nature07484
发表时间: 2008-11-06
期刊: NATURE
影响因子: 64.8
作者:
Wang, Jun;Wang, Wei;Li, Ruiqiang;Li, Yingrui;Tian, Geng;Goodman, Laurie;Fan, Wei;Zhang, Junqing;Li, Jun;Zhang, Juanbin;Guo, Yiran;Feng, Binxiao;Li, Heng;Lu, Yao;Fang, Xiaodong;Liang, Huiqing;Du, Zhenglin;Li, Dong;Zhao, Yiqing;Hu, Yujie;Yang, Zhenzhen;Zheng, Hancheng;Hellmann, Ines;Inouye, Michael;Pool, John;Yi, Xin;Zhao, Jing;Duan, Jinjie;Zhou, Yan;Qin, Junjie;Ma, Lijia;Li, Guoqing;Yang, Zhentao;Zhang, Guojie;Yang, Bin;Yu, Chang;Liang, Fang;Li, Wenjie;Li, Shaochuan;Li, Dawei;Ni, Peixiang;Ruan, Jue;Li, Qibin;Zhu, Hongmei;Liu, Dongyuan;Lu, Zhike;Li, Ning;Guo, Guangwu;Zhang, Jianguo;Ye, Jia;Fang, Lin;Hao, Qin;Chen, Quan;Liang, Yu;Su, Yeyang;San, A.;Ping, Cuo;Yang, Shuang;Chen, Fang;Li, Li;Zhou, Ke;Zheng, Hongkun;Ren, Yuanyuan;Yang, Ling;Gao, Yang;Yang, Guohua;Li, Zhuo;Feng, Xiaoli;Kristiansen, Karsten;Wong, Gane Ka-Shu;Nielsen, Rasmus;Durbin, Richard;Bolund, Lars;Zhang, Xiuqing;Li, Songgang;Yang, Huanming;Wang, Jian
通讯作者: Wang, Jian
DOI: 10.1186/gb-2009-10-3-r25
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Langmead B;Trapnell C;Pop M;Salzberg SL
通讯作者: Salzberg SL
DOI: 10.1145/1327452.1327492
发表时间: 2008-01-01
影响因子: 22.7
作者:
Dean, Jeffrey;Ghemawat, Sanjay
通讯作者: Ghemawat, Sanjay