Sieve: Scalable In-situ DRAM-based Accelerator Designs for Massively Parallel k-mer Matching

Sieve: Scalable In-situ DRAM-based Accelerator Designs for Massively Parallel k-mer Matching
复制标题

DOI:
10.1109/isca52012.2021.00028
复制
发表时间:
2021-06
期刊:
2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Lingxi Wu;Rasool Sharifi;Marzieh Lenjani;K. Skadron;A. Venkat
Lingxi Wu;Rasool Sharifi;Marzieh Lenjani;K. Skadron;A. Venkat
中科院分区:
其他
文献类型:
--
作者:
Lingxi Wu;Rasool Sharifi;Marzieh Lenjani;K. Skadron;A. Venkat

文献摘要

相似文献

Biosequence数据的迅速涌入,再加上现代计算系统的处理能力的停滞,突出了探索探索高性能加速器的关键需求,以满足现代生物信息信息应用的不断增加的吞吐量需求。这项工作认为,内存中的处理(PIM)是提高K-MER匹配性能的有效解决方案,K-MER匹配是标准生物信息学管道中关键的瓶颈阶段,其特征是随机访问模式和低计算强度。这项工作提出了三个DRAM - 基于原位K-MER匹配的加速器设计(一种针对区域进行了优化,一个针对吞吐量进行了优化的设计,并且在硬件成本和性能之间达到平衡),称为筛子,该筛子利用新颖的数据映射方案可以同时进行比较数以百万计的DNA碱基对,用于快速图案匹配的轻质匹配电路以及一种早期终止机制,该机制可预防不必要的DRAM行激活,以减少潜伏期并节省能量。使用现实世界中数据集对筛网的评估表明,在多核CPU/GPU的基地上,最具侵略性的设计平均提供了326X/32X的加速和74x/48x的能源节省匹配。
The rapid influx of biosequence data, coupled with the stagnation of the processing power of modern computing systems, highlights the critical need for exploring high-performance accelerators that can meet the ever-increasing throughput demands of modern bioinformatics applications. This work argues that processing in memory (PIM) is an effective solution to enhance the performance of k-mer matching, a critical bottleneck stage in standard bioinformatics pipelines, that is characterized by random access patterns and low computational intensity.This work proposes three DRAM-based in-situ k-mer matching accelerator designs (one optimized for area, one optimized for throughput, and one that strikes a balance between hardware cost and performance), dubbed Sieve, that leverage a novel data mapping scheme to allow for simultaneous comparisons of millions of DNA base pairs, lightweight matching circuitry for fast pattern matching, and an early termination mechanism that prunes unnecessary DRAM row activation to reduce latency and save energy. Evaluation of Sieve using state-of-the-art workloads with real-world datasets shows that the most aggressive design provides an average of 326x/32x speedup and 74X/48x energy savings over multi-core-CPU/GPU baselines for k-mer matching.