HIRAC: A Hierarchical Accelerator with Sorting-based Packing for SpGEMMs in DNN Applications

HIRAC: A Hierarchical Accelerator with Sorting-based Packing for SpGEMMs in DNN Applications
复制标题

DOI:
10.1109/hpca56546.2023.10070977
复制
发表时间:
2023-02
期刊:
2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Hesam Shabani;Abhishek Singh;Bishoy Youhana;Xiaochen Guo
Hesam Shabani;Abhishek Singh;Bishoy Youhana;Xiaochen Guo
中科院分区:
其他
文献类型:
--
作者:
Hesam Shabani;Abhishek Singh;Bishoy Youhana;Xiaochen Guo

文献摘要

相似文献

最先进的深度神经网络(DNN)模型使用修剪来避免过拟合并减少参数数量。为了提高存储和计算效率,只存储非零元素,并将其位置编码为稀疏格式。稀疏通用矩阵乘法(SpGEMM)是基于深度神经网络应用的核心计算。计算SpGEMM的一个挑战是,在由处理元素(PE)阵列组成的硬件加速器中,要避免将零元素相乘,同时保持较高的硬件利用率。先前解决这一挑战的工作通常需要复杂的互连网络,这增加了高面积和能源成本。这项工作提出了一种硬件/软件协同设计架构,可以在不需要复杂互连网络的情况下高效地计算SpGEMM。提出了一种新的快速打包算法SorPack,将稀疏矩阵转换为密集矩阵,提高了PE利用率。关键思想是根据非零元素的数量对每个子矩阵中的列和行进行排序。目标是保持需要加在一起的部分和彼此接近,因此可以在本地加,避免使用复杂的互连网络。此外,提出了一种新的基于tile的分层体系结构HIRAC,以提供一个可扩展的系统,最大限度地提高pe的并行性。HIRAC架构由新颖的PE阵列设计和为DNN应用量身定制的互连网络组成。SorPack算法补充了HIRAC,进一步提高了硬件利用率和整体系统性能。根据评估结果,与最先进的稀疏DNN加速器SIGMA相比,HIRAC在单层DNN上实现了平均3.2倍的加速。此外,与SIGMA相比,HIRAC的面积减少了9.5%,功耗降低了32%。对DNN模型的端到端评估显示,在TPU上运行时间减少了8.2倍。
The state-of-the-art deep neural network (DNN) models use pruning to avoid over-fitting and reduce the number of parameters. In order to improve storage and computational efficiency, only nonzero elements are stored, and their locations are encoded into a sparse format. Sparse General Matrix Multiplication (SpGEMM) is the kernel computation of DNN-based applications. One challenge of computing SpGEMM is to avoid multiplying zero elements while keeping hardware utilization high in hardware accelerators that consist of processing element (PE) arrays. Prior work tackling this challenge typically requires complex interconnection networks, which adds high area and energy costs.This work proposes a HW/SW co-design architecture to compute SpGEMM efficiently without requiring complex interconnection networks. A novel fast packing algorithm, SorPack, is proposed to convert a sparse matrix into a dense matrix that increases PE utilization. The key idea is to sort columns and rows inside each submatrix based on the number of nonzero elements. The goal is to keep the partial sums that need to be added together close to each other, hence can be added locally and avoid the use of complex interconnection networks. In addition, a new tile-based hierarchical architecture, HIRAC, is proposed to provide a scalable system that maximizes the parallelism of the PEs. The HIRAC architecture consists of a novel PE array design and interconnection network tailored for DNN applications. The SorPack algorithm complements the HIRAC to further improve hardware utilization and overall system performance. Based on the evaluation results, HIRAC achieves an average of 3.2× speedup on a single layer of DNN as compared to the state-of-the-art sparse DNN accelerator SIGMA. In addition, HIRAC has a 9.5% area reduction and a 32% power reduction as compared to SIGMA. An end-to-end evaluation on a DNN model shows an 8.2× runtime reduction over the TPU.