EPA-ng: Massively Parallel Evolutionary Placement of Genetic Sequences.

EPA-ng: Massively Parallel Evolutionary Placement of Genetic Sequences.
复制标题

EPA-NG:遗传序列的大量平行进化位置。

DOI:
10.1093/sysbio/syy054
复制
发表时间:
2019-03-01
期刊:
影响因子:
6.5
通讯作者:
Stamatakis A
Stamatakis A
中科院分区:
生物学1区
文献类型:
--
作者:
Barbera P;Kozlov AM;Czech L;Morel B;Darriba D;Flouri T;Stamatakis A

文献摘要

参考文献

被引文献

相似文献

下一代测序(NGS)技术带来了无处不在的分子序列数据。这种数据雪崩在元遗传学中尤其具有挑战性,它侧重于对从不同微生物环境中获得的序列进行分类鉴定。系统发育定位方法决定了这些序列如何适应进化背景。以前的系统发育布局算法的实现,如RAxML中包含的进化布局算法(EPA),或PPLACER,正越来越多地用于这一目的。然而,由于NGS技术的稳步发展,当前的实现面临着相当大的可扩展性限制。在这里,我们提出了EPA-NG,它是EPA的完全重新实现,速度大大加快,提供了分布式内存并行化,并集成了RAxML-EPA和PPLACER的概念。EPA-NG可以在标准共享内存上执行,也可以在分布式内存系统(例如计算集群)上执行。为了展示EPA-NG的可扩展性,我们将来自Tara Ocean项目的10亿元遗传读数放在一棵有3748个分类群的参考树上,使用了2048个核心。我们的性能评估表明,在顺序执行模式下,EPA-NG的性能比RAxML-EPA和PPLACER高出一倍,同时在共享存储系统上获得了相当的并行效率。我们进一步证明,EPA-NG的分布式存储并行化可以很好地扩展到2048核。环保局-NG在AGPLv3许可证下可用:https://github.com/Pbdas/epa-ng.
Next generation sequencing (NGS) technologies have led to a ubiquity of molecular sequence data. This data avalanche is particularly challenging in metagenetics, which focuses on taxonomic identification of sequences obtained from diverse microbial environments. Phylogenetic placement methods determine how these sequences fit into an evolutionary context. Previous implementations of phylogenetic placement algorithms, such as the evolutionary placement algorithm (EPA) included in RAxML, or PPLACER, are being increasingly used for this purpose. However, due to the steady progress in NGS technologies, the current implementations face substantial scalability limitations. Herein, we present EPA-NG, a complete reimplementation of the EPA that is substantially faster, offers a distributed memory parallelization, and integrates concepts from both, RAxML-EPA and PPLACER. EPA-NG can be executed on standard shared memory, as well as on distributed memory systems (e.g., computing clusters). To demonstrate the scalability of EPA-NG, we placed billion metagenetic reads from the Tara Oceans Project onto a reference tree with 3748 taxa in just under h, using 2048 cores. Our performance assessment shows that EPA-NG outperforms RAxML-EPA and PPLACER by up to a factor of in sequential execution mode, while attaining comparable parallel efficiency on shared memory systems. We further show that the distributed memory parallelization of EPA-NG scales well up to 2048 cores. EPA-NG is available under the AGPLv3 license: https://github.com/Pbdas/epa-ng.
DOI: 10.1126/science.1261359
发表时间: 2015-05-22
期刊: SCIENCE
影响因子: 56.9
作者:
Sunagawa, Shinichi;Coelho, Luis Pedro;Bork, Peer
通讯作者: Bork, Peer
DOI: 10.1038/s41559-017-00911
发表时间: 2017-04-01
影响因子: 16.8
作者:
Mahe, Frederic;de Vargas, Colomban;Dunthorn, Micah
通讯作者: Dunthorn, Micah
DOI: 10.1038/nature12171
发表时间: 2013-06-20
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1093/nar/gkw396
发表时间: 2016-06-20
影响因子: 14.9
作者:
Kozlov, Alexey M.;Zhang, Jiajie;Stamatakis, Alexandros
通讯作者: Stamatakis, Alexandros
DOI: 10.1111/j.1467-9868.2011.01018.x
发表时间: 2012-06-01
期刊: Journal of the Royal Statistical Society. Series B, Statistical methodology
影响因子: --
作者:
Evans SN;Matsen FA
通讯作者: Matsen FA