ADEPT: a domain independent sequence alignment strategy for gpu architectures

ADEPT: a domain independent sequence alignment strategy for gpu architectures
复制标题

DOI:
10.1186/s12859-020-03720-1
复制
发表时间:
2020-09-15
期刊:
影响因子:
3
通讯作者:
Yelick, Katherine
Yelick, Katherine
中科院分区:
生物学4区
文献类型:
--
作者:
Awan, Muaaz G.;Deslippe, Jack;Yelick, Katherine

文献摘要

被引文献

相似文献

背景生物信息学工作流程经常使用自动化基因组组装和蛋白质聚类工具。在大多数这些工具的核心中,执行时间的很大一部分花费在确定两个序列之间的最佳局部比对上。这个任务是用史密斯-沃特曼算法,这是一个基于动态规划的方法。随着现代测序技术的出现以及基因组和蛋白质数据库规模的增加,出现了对更快的Smith-Waterman实现的需求。多个SIMD策略的史密斯-沃特曼算法可用于CPU。然而,随着HPC设施向基于加速器的架构的转变,出现了对高效GPU加速策略的需求。现有的基于GPU的策略要么针对特定类型的字符(核苷酸或氨基酸)进行了优化,要么仅针对少数应用程序用例进行了优化。结果在本文中,我们提出了ADEPT,一个新的序列比对策略的GPU架构,是域独立的,支持从基因组和蛋白质序列的比对。我们提出的策略使用GPU特定的优化,不依赖于序列的性质。我们通过实施史密斯-沃特曼算法并将其与类似的CPU策略以及每个域的最快已知GPU方法进行比较来证明这种策略的可行性。ADEPT的驱动程序使其能够跨多个GPU扩展,并允许轻松集成到利用大规模计算系统的软件管道中。我们已经表明,基于ADEPT的Smith-Waterman算法在科里超级计算机的单个GPU节点(8个GPU)上分别对基于蛋白质和基于DNA的数据集表现出360 GCUPS和497 GCUPS的峰值性能。总体而言,ADEPT在节点到节点的比较中显示出比相应的SIMD CPU实现快10倍的性能。结论ADEPT展示了与现有GPU策略相当或更好的性能。我们通过将ADEPT集成到MetaHipMer(一种高性能从头宏基因组组装程序)和PASTIS(一种高性能蛋白质相似性图构建管道)中,证明了ADEPT在支持现有生物信息学软件管道方面的有效性。我们的结果显示,MetaHipMer和PASTIS的性能分别提高了10%和30%。
Background Bioinformatic workflows frequently make use of automated genome assembly and protein clustering tools. At the core of most of these tools, a significant portion of execution time is spent in determining optimal local alignment between two sequences. This task is performed with the Smith-Waterman algorithm, which is a dynamic programming based method. With the advent of modern sequencing technologies and increasing size of both genome and protein databases, a need for faster Smith-Waterman implementations has emerged. Multiple SIMD strategies for the Smith-Waterman algorithm are available for CPUs. However, with the move of HPC facilities towards accelerator based architectures, a need for an efficient GPU accelerated strategy has emerged. Existing GPU based strategies have either been optimized for a specific type of characters (Nucleotides or Amino Acids) or for only a handful of application use-cases. Results In this paper, we present ADEPT, a new sequence alignment strategy for GPU architectures that is domain independent, supporting alignment of sequences from both genomes and proteins. Our proposed strategy uses GPU specific optimizations that do not rely on the nature of sequence. We demonstrate the feasibility of this strategy by implementing the Smith-Waterman algorithm and comparing it to similar CPU strategies as well as the fastest known GPU methods for each domain. ADEPT's driver enables it to scale across multiple GPUs and allows easy integration into software pipelines which utilize large scale computational systems. We have shown that the ADEPT based Smith-Waterman algorithm demonstrates a peak performance of 360 GCUPS and 497 GCUPs for protein based and DNA based datasets respectively on a single GPU node (8 GPUs) of the Cori Supercomputer. Overall ADEPT shows 10x faster performance in a node-to-node comparison against a corresponding SIMD CPU implementation. Conclusions ADEPT demonstrates a performance that is either comparable or better than existing GPU strategies. We demonstrated the efficacy of ADEPT in supporting existing bionformatics software pipelines by integrating ADEPT in MetaHipMer a high-performance denovo metagenome assembler and PASTIS a high-performance protein similarity graph construction pipeline. Our results show 10% and 30% boost of performance in MetaHipMer and PASTIS respectively.