HARVEST: Towards Efficient Sparse DNN Accelerators using Programmable Thresholds

HARVEST: Towards Efficient Sparse DNN Accelerators using Programmable Thresholds
复制标题

收获:使用可编程阈值实现高效稀疏 DNN 加速器

DOI:
10.1109/vlsid60093.2024.00044
复制
发表时间:
2024
期刊:
2024 37th International Conference on VLSI Design and 2024 23rd International Conference on Embedded Systems (VLSID)
影响因子:
--
通讯作者:
Vijay Raghunathan
Vijay Raghunathan
中科院分区:
--
文献类型:
--
作者:
Soumendu Kumar Ghosh;Shamik Kundu;Arnab Raha;Deepak A. Mathaikutty;Vijay Raghunathan

文献摘要

被引文献

相似文献

尽管由于算法的进步,DNN已成为人工智能标准,但其计算和内存需求对其在边缘设备上的部署构成了挑战。广泛的研究探索了在资源受限的移动的设备上有效运行DNN模型的优化,包括软件和硬件增强。稀疏性利用是一种重要的优化技术,旨在通过消除零操作数导致的冗余MAC操作来提高DNN推理效率和速度。在本文中,我们提出了HARVEST,这是一种硬件-软件协同设计方法,它利用加速器中现有的稀疏引擎在DNN权重和推理过程中的激活中引入稀疏性。该技术包括两种方法:(1)通过对中间激活应用阈值来实现激活稀疏性。这些阈值是基于存储在加速器的SRAM组中的激活的约束和统计来确定的。这种方法将阈值化应用于所有层类型,包括卷积层、逐元素层、全连接层、注意层、归一化层和非线性激活函数,从而增加稀疏性,而不会在面积和功耗方面产生任何额外开销。(2)权重稀疏性在部署之前使用每个层的自定义阈值引入。这些方法共同降低了内存和计算能耗,从而提高了加速器的能效,同时以最小的精度损失为代价。最先进的Transformer和CNN模型的结果表明,激活和权重稀疏性分别提高了61%和80%。利用这种稀疏性,内部完全稀疏的加速器分别在内存、计算和总体加速器能量方面减少了24%、36%和32%,而准确性损失最小(< 0.5%)。此外,HARVEST在DNN推理期间提供高达32%和36%的内存和计算周期计数减少,从而提高吞吐量。
Although DNNs have become the AI standard due to algorithmic advancements, their computational and memory demands pose challenges for their deployment on edge devices. Extensive research has explored optimizations to efficiently run DNN models on resource-constrained mobile devices, encompassing both software and hardware enhancements. Sparsity exploitation is a prominent optimization technique that aims to boost DNN inference efficiency and speed by eliminating redundant MAC operations resulting from zero operands. In this paper, we propose HARVEST, a hardware-software co-design approach that utilizes existing sparsity engines in accelerators to introduce sparsity in DNN weights and activations during inference. This technique involves two methods: (1) Activation sparsity is achieved by applying thresholds to intermediate activations. These thresholds are determined based on constraints and statistics of the activations stored in the accelerator’s SRAM banks. This approach applies thresholding to all layer types, including convolution, element-wise, fully connected, attention, normalization layers, and non-linear activation functions, thereby increasing sparsity without any additional overhead in terms of area and power. (2) Weight sparsity is introduced before deployment using customized thresholds for each layer. These methods collectively reduce memory and compute energy consumption, leading to improvements in accelerator energy efficiency, at the cost of minimal accuracy loss. Results on state-of-the-art Transformer and CNN models demonstrate a gain of up to 61% and 80% in activation and weight sparsity, respectively. Exploiting this sparsity, an in-house fully-sparse accelerator provides up to 24%, 36%, and 32% reductions in memory, compute, and overall accelerator energy, respectively, for minimal (< 0.5%) loss in accuracy. Furthermore, HARVEST provides up to 32% and 36% reduction in memory and compute cycle count during DNN inference, leading to increased throughput.