MSAProbs-MPI: parallel multiple sequence aligner for distributed-memory systems

MSAProbs-MPI: parallel multiple sequence aligner for distributed-memory systems
复制标题

DOI:
10.1093/bioinformatics/btw558
复制
发表时间:
2016-12-15
期刊:
影响因子:
5.8
通讯作者:
Schmidt, Bertil
Schmidt, Bertil
中科院分区:
生物学3区
文献类型:
--
作者:
Gonzalez-Dominguez, Jorge;Liu, Yongchao;Schmidt, Bertil

文献摘要

被引文献

相似文献

MSAProbs是一种基于隐马尔可夫模型的最先进的蛋白质多序列比对工具。对于大规模的输入数据集,它可以以相对较长的运行时间为代价来实现高对齐精度。在这项工作中,我们提出了MSAProbs-MPI,它是多线程MSAProbs工具的分布式内存并行版本,能够通过利用常见多核CPU集群的计算能力来减少运行时间。我们在具有32个节点(每个节点包含两个Intel Haswell处理器)的集群上进行的性能评估显示,对于典型的输入数据集,执行时间减少了一个数量级以上。此外,使用8个节点的MSAProbs-MPI比运行在特斯拉K20上的GPU加速的QuickProbs更快。另一个优点是,MSAProbs-MPI可以处理MSAProbs和QuickProbs可能分别由于时间和内存限制而失败的大型数据集。可用性和实施:在LINUX系统上运行的C++和MPI源代码以及参考手册,请访问http://msaprobs.sourceforge.netContact:jGonzalezd@udc.es.补充信息:补充数据可在BioInformation Online上找到。
MSAProbs is a state-of-the-art protein multiple sequence alignment tool based on hidden Markov models. It can achieve high alignment accuracy at the expense of relatively long runtimes for large-scale input datasets. In this work we present MSAProbs-MPI, a distributed-memory parallel version of the multithreaded MSAProbs tool that is able to reduce runtimes by exploiting the compute capabilities of common multicore CPU clusters. Our performance evaluation on a cluster with 32 nodes (each containing two Intel Haswell processors) shows reductions in execution time of over one order of magnitude for typical input datasets. Furthermore, MSAProbs-MPI using eight nodes is faster than the GPU-accelerated QuickProbs running on a Tesla K20. Another strong point is that MSAProbs-MPI can deal with large datasets for which MSAProbs and QuickProbs might fail due to time and memory constraints, respectively.Availability and Implementation: Source code in C++ and MPI running on Linux systems as well as a reference manual are available at http://msaprobs.sourceforge.netContact: jgonzalezd@udc.esSupplementary information: Supplementary data are available at Bioinformatics online.