TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for Transformer

TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for Transformer
复制标题

DOI:
10.1109/hpca53966.2022.00082
复制
发表时间:
2022-04
期刊:
2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Minxuan Zhou;Weihong Xu;Jaeyoung Kang;Tajana Simunic
Minxuan Zhou;Weihong Xu;Jaeyoung Kang;Tajana Simunic
中科院分区:
其他
文献类型:
--
作者:
Minxuan Zhou;Weihong Xu;Jaeyoung Kang;Tajana Simunic

文献摘要

被引文献

相似文献

基于transformer的模型是许多机器学习(ML)任务的最新技术。执行Transformer通常需要很长的执行时间,因为内存占用量大,数据重用率低,给内存系统带来压力,同时计算资源利用不足。基于内存的处理技术,包括内存内处理(PIM)和近内存计算(NMC),有望加速Transformer,因为它们提供了高的内存带宽利用率和广泛的计算并行性。然而,之前基于内存的ML加速器主要针对计算密集型ML模型(例如,CNN),这不符合Transformer的内存密集型特性。在这项工作中,我们提出了transPIM,一个基于内存的加速Transformer使用软件和硬件协同设计。在软件层,TransPIM采用了基于令牌的层间数据流,避免了以前的基于层的层间数据流所引入的昂贵的层间数据移动。在硬件层面,TransPIM在传统的高带宽存储器(HBM)架构中引入了轻量级修改,以支持PIM-NMC混合处理和高效的数据通信,从而加速基于Transformer的模型。我们的实验表明,TransPIM比现有的基于内存的加速快3.7倍到9.1倍。与传统加速器相比,TransPIM比GPU快22.1倍至114.9倍,并且提供比现有基于ASIC的加速器高2.0倍的吞吐量。
Transformer-based models are state-of-the-art for many machine learning (ML) tasks. Executing Transformer usually requires a long execution time due to the large memory footprint and the low data reuse rate, stressing the memory system while under-utilizing the computing resources. Memory-based processing technologies, including processing in-memory (PIM) and near-memory computing (NMC), are promising to accelerate Transformer since they provide high memory bandwidth utilization and extensive computation parallelism. However, the previous memory-based ML accelerators mainly target at optimizing dataflow and hardware for compute-intensive ML models (e.g., CNNs), which do not fit the memory-intensive characteristics of Transformer. In this work, we propose TransPIM, a memory-based acceleration for Transformer using software and hardware co-design. In the software-level, TransPIM adopts a token-based dataflow to avoid the expensive inter-layer data movements introduced by previous layer-based dataflow. In the hardware-level, TransPIM introduces lightweight modifications in the conventional high bandwidth memory (HBM) architecture to support PIM-NMC hybrid processing and efficient data communication for accelerating Transformer-based models. Our experiments show that TransPIM is 3.7× to 9.1× faster than existing memory-based acceleration. As compared to conventional accelerators, TransPIM is 22.1× to 114.9× faster than GPUs and provides 2.0× more throughput than existing ASIC-based accelerators.