CANDLES: Channel-Aware Novel Dataflow-Microarchitecture Co-Design for Low Energy Sparse Neural Network Acceleration

CANDLES: Channel-Aware Novel Dataflow-Microarchitecture Co-Design for Low Energy Sparse Neural Network Acceleration
复制标题

DOI:
10.1109/hpca53966.2022.00069
复制
发表时间:
2022-04
期刊:
2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Sumanth Gudaparthi;Sarabjeet Singh;Surya Narayanan;R. Balasubramonian;V. Sathe
Sumanth Gudaparthi;Sarabjeet Singh;Surya Narayanan;R. Balasubramonian;V. Sathe
中科院分区:
其他
文献类型:
--
作者:
Sumanth Gudaparthi;Sarabjeet Singh;Surya Narayanan;R. Balasubramonian;V. Sathe

文献摘要

被引文献

相似文献

为了利用深度神经网络激活和权值所表现出的稀疏性,已经设计了几种深度神经网络(DNN)加速器。最先进的稀疏加速器可以被描述为像素优先或通道优先加速器,每个加速器都有其独特的数据流和压缩格式来帮助其数据流。前者在更新神经元部分和上花费大量的能量,而后者在处理索引元数据上花费大量的能量。这项工作引入了一种新的微架构和数据流,通过采用像素优先压缩和通道优先数据流来协调这些权衡。所提出的微体系结构具有更简单的索引生成逻辑,并结合了累加器缓冲区层次结构和具有低布线开销的交叉栏。压缩格式和数据流促进了神经元更新的高时间局部性,进一步降低了能量。最后,我们引入跨处理元素的工作分区,这自然会导致负载平衡,而无需离线分析。与四个最先进的基线相比,拟议的建筑,蜡烛,明显优于三个,并与第四个的性能相匹配。在能源方面,candle比这四个基准节能2.5 - 5.6倍。
Several deep neural network (DNN) accelerators have been designed to exploit the sparsity exhibited by DNN activations and weights. State-of-the-art sparse accelerators can be described as either Pixel-first or Channel-first accelerators, each with its unique dataflow and compression format aiding its dataflow. The former expends significant energy updating neuron partial sums, while the latter expends significant energy in handling the index metadata. This work introduces a novel microarchitecture and dataflow that reconciles these trade-offs by adopting a Pixel-first compression and Channel-first dataflow. The proposed microarchitecture has a simpler index-generation logic combined with an accumulator buffer hierarchy and crossbar with low wiring overhead. The compression format and dataflow promote high temporal locality in neuron updates, further lowering energy. Finally, we introduce work partitions across processing elements that naturally lead to load balance without offline analysis. Compared to four state-of-the-art baselines, the proposed architecture, CANDLES, significantly outperforms three and matches the performance of the fourth. In terms of energy, CANDLES is between 2.5× and 5.6× more energy-efficient than these four baselines.