MetaSpark: a spark-based distributed processing tool to recruit metagenomic reads to reference genomes

MetaSpark: a spark-based distributed processing tool to recruit metagenomic reads to reference genomes
复制标题

MetaSpark:基于 Spark 的分布式处理工具,用于将宏基因组读数招募到参考基因组

DOI:
10.1093/bioinformatics/btw750
复制
发表时间:
2017
期刊:
影响因子:
5.8
通讯作者:
Niu Beifang
Niu Beifang
中科院分区:
生物学3区
文献类型:
--
作者:
Zhou Wei;Li Ruilin;Yuan Shuo;Liu ChangChun;Yao Shaowen;Luo Jing;Niu Beifang

文献摘要

相似文献

摘要随着下一代测序的到来,传统的生物信息学工具受到了海量原始元基因组数据的挑战。元基因组研究的瓶颈之一是缺乏适合大规模和云计算的数据分析工具。在本文中,我们提出了一个基于Spark的工具,称为MetaSpark,用于招募元基因组阅读到参考基因组。MetaSpark得益于Spark的分布式数据集(RDD),这使得它能够跨集群节点缓存内存中的数据集,并随数据集进行良好的扩展。与之前的元基因组学招募工具相比,MetaSpark招募的阅读量明显高于SOAP2、BWA和LAST等许多程序,并且与FR-HIT相比,当有100万个阅读量和0.75Gb引用时,MetaSpark招募的阅读量增加了∼4%。不同的测试案例展示了MetaSpark的可扩展性和总体较高的performance.Availabilityhttps://github.com/zhouweiyg/metasparkSupplementary信息补充数据可在BioInformation Online上获得
SummaryWith the advent of next-generation sequencing, traditional bioinformatics tools are challenged by massive raw metagenomic datasets. One of the bottlenecks of metagenomic studies is lack of large-scale and cloud computing suitable data analysis tools. In this paper, we proposed a Spark-based tool, called MetaSpark, to recruit metagenomic reads to reference genomes. MetaSpark benefits from the distributed data set (RDD) of Spark, which makes it able to cache data set in memory across cluster nodes and scale well with the datasets. Compared with previous metagenomics recruitment tools, MetaSpark recruited significantly more reads than many programs such as SOAP2, BWA and LAST and increased recruited reads by ∼4% compared with FR-HIT when there were 1 million reads and 0.75 GB references. Different test cases demonstrate MetaSpark’s scalability and overall high performance.Availabilityhttps://github.com/zhouweiyg/metasparkSupplementary informationSupplementary data are available atBioinformaticsonline