An overview of the Hadoop/MapReduce/HBase framework and its current applications in bioinformatics.

An overview of the Hadoop/MapReduce/HBase framework and its current applications in bioinformatics.
复制标题

DOI:
10.1186/1471-2105-11-s12-s1
复制
发表时间:
2010-12-21
期刊:
影响因子:
3
通讯作者:
Taylor RC
Taylor RC
中科院分区:
生物学4区
文献类型:
--
作者:
Taylor RC

文献摘要

被引文献

相似文献

生物信息学研究人员现在面临着超大规模数据集的分析,这个问题在未来几年只会以惊人的速度增长。开源软件的最新发展,即Hadoop项目和相关软件,为扩展到Linux集群上的pb级数据仓库提供了基础,使用名为MapReduce的编程风格提供了对此类数据的容错并行分析。概述了Hadoop(顶级Apache软件基金会项目)和相关开源软件项目在生物信息学社区中的当前使用情况。定义了Hadoop及其相关的HBase项目背后的概念,并描述了当前使用Hadoop的生物信息学软件。重点是下一代测序,作为迄今为止领先的应用领域。Hadoop和MapReduce编程范式已经在生物信息学社区,特别是在下一代测序分析领域有了坚实的基础,而且这种使用正在增加。这是由于基于Hadoop的分析在商品Linux集群上的成本效益,以及在云端通过数据上传到已经实现Hadoop/HBase的云供应商;并且由于MapReduce方法在并行化许多数据分析算法方面的有效性和易用性。
Bioinformatics researchers are now confronted with analysis of ultra large-scale data sets, a problem that will only increase at an alarming rate in coming years. Recent developments in open source software, that is, the Hadoop project and associated software, provide a foundation for scaling to petabyte scale data warehouses on Linux clusters, providing fault-tolerant parallelized analysis on such data using a programming style named MapReduce. An overview is given of the current usage within the bioinformatics community of Hadoop, a top-level Apache Software Foundation project, and of associated open source software projects. The concepts behind Hadoop and the associated HBase project are defined, and current bioinformatics software that employ Hadoop is described. The focus is on next-generation sequencing, as the leading application area to date. Hadoop and the MapReduce programming paradigm already have a substantial base in the bioinformatics community, especially in the field of next-generation sequencing analysis, and such use is increasing. This is due to the cost-effectiveness of Hadoop-based analysis on commodity Linux clusters, and in the cloud via data upload to cloud vendors who have implemented Hadoop/HBase; and due to the effectiveness and ease-of-use of the MapReduce method in parallelization of many data analysis algorithms.