cuBLASTP: Fine-Grained Parallelization of Protein Sequence Search on CPU+GPU

cuBLASTP: Fine-Grained Parallelization of Protein Sequence Search on CPU+GPU
复制标题

cuBLASTP:CPU GPU 上蛋白质序列搜索的细粒度并行化

DOI:
--
复制
发表时间:
2014
期刊:
IEEE/ACM Transactions on Computational Biology & Bioinformatics
影响因子:
--
通讯作者:
Wu
Wu
中科院分区:
--
文献类型:
--
作者:
Jing Zhang;Hao Wang;Heshan Lin;Wu

文献摘要

被引文献

相似文献

BLAST是基本局部比对搜索工具的缩写,是在生命科学中用于成对序列搜索的普遍存在的工具。然而,随着下一代测序(NGS)的出现,无论是在NGS的开始还是下游,序列数据库的指数增长都超过了我们分析数据的能力。虽然最近的研究已经利用图形处理单元(GPU)来加速用于搜索蛋白质序列的BLAST算法(即,BLASTP),这些研究使用粗粒度的并行性,其中一个序列比对仅映射到一个线程。这种方法不能有效地利用GPU的能力,特别是由于BLASTP在执行路径和存储器访问模式两者中的不规则性。为了解决上述缺点,我们提出了一种细粒度的方法来并行化BLASTP,其中每个单独的序列搜索阶段被映射到GPU上的许多线程。这种方法,我们称之为cuBLASTP,重新排序数据访问模式,并减少最耗时阶段的发散分支(即,命中检测和无空位扩展)。此外,cuBLASTP优化了剩余的相(即,有间隙的扩展和与回溯的对齐),并将它们的执行与GPU上运行的阶段重叠。
BLAST, short for Basic Local Alignment Search Tool, is a ubiquitous tool used in the life sciences for pairwise sequence search. However, with the advent of next-generation sequencing (NGS), whether at the outset or downstream from NGS, the exponential growth of sequence databases is outstripping our ability to analyze the data. While recent studies have utilized the graphics processing unit (GPU) to speedup the BLAST algorithm for searching protein sequences (i.e., BLASTP), these studies use coarse-grained parallelism, where one sequence alignment is mapped to only one thread. Such an approach does not efficiently utilize the capabilities of a GPU, particularly due to the irregularity of BLASTP in both execution paths and memory-access patterns. To address the above shortcomings, we present a fine-grained approach to parallelize BLASTP, where each individual phase of sequence search is mapped to many threads on a GPU. This approach, which we refer to as cuBLASTP, reorders data-access patterns and reduces divergent branches of the most time-consuming phases (i.e., hit detection and ungapped extension). In addition, cuBLASTP optimizes the remaining phases (i.e., gapped extension and alignment with trace back) on a multicore CPU and overlaps their execution with the phases running on the GPU.