Libra: Improved Partitioning Strategies for Massive Comparative Metagenomics Analysis

Libra: Improved Partitioning Strategies for Massive Comparative Metagenomics Analysis
复制标题

Libra:改进的大规模比较宏基因组分析的分区策略

DOI:
10.1145/3217880.3217882
复制
发表时间:
2018
期刊:
Proceedings of the 9th Workshop on Scientific Cloud Computing
影响因子:
--
通讯作者:
Hartman, John H.
Hartman, John H.
中科院分区:
--
文献类型:
--
作者:
Choi, Illyoung;Ponsero, Alise J.;Youens-Clark, Ken;Bomhoff, Matthew;Hurwitz, Bonnie L.;Hartman, John H.

文献摘要

参考文献

相似文献

大数据分析平台(例如 Hadoop)对科学计算很有吸引力,因为它们无处不在、得到良好支持且易于理解。不幸的是,负载平衡是在这些平台上实现大规模科学计算应用程序的常见挑战。在本文中,我们介绍了 Libra 的设计和实现,这是一种基于 Hadoop 的比较宏基因组学工具(比较从环境中收集的遗传物质样本)。我们描述了 Libra 执行的计算以及如何使用 Hadoop 任务实现该计算,包括 Libra 使用的技术来确保任务工作负载在样本大小不均匀和样本中遗传物质分布不均的情况下保持平衡。在 10 台机器的 Hadoop 集群上,Libra 可以在不到 20 小时内分析整个 Tara Ocean Virome 的约 42 亿个读取。
Big-data analytics platforms, such as Hadoop, are appealing for scientific computation because they are ubiquitous, well-supported, and well-understood. Unfortunately, load-balancing is a common challenge of implementing large-scale scientific computing applications on these platforms. In this paper we present the design and implementation of Libra, a Hadoop-based tool for comparative metagenomics (comparing samples of genetic material collected from the environment). We describe the computation that Libra performs and how that computation is implemented using Hadoop tasks, including the techniques used by Libra to ensure that the task workloads are balanced despite nonuniform sample sizes and skewed distributions of genetic material in the samples. On a 10-machine Hadoop cluster Libra can analyze the entire Tara Ocean Viromes of ~4.2 billion reads in fewer than 20 hours.
DOI: 10.1109/tcbb.2017.2760829
发表时间: 2019-07-01
影响因子: 4.5
作者:
Pan,Tony;Flick,Patrick;Aluru,Srinivas
通讯作者: Aluru,Srinivas