ESCALATE: Boosting the Efficiency of Sparse CNN Accelerator with Kernel Decomposition

ESCALATE: Boosting the Efficiency of Sparse CNN Accelerator with Kernel Decomposition
复制标题

DOI:
10.1145/3466752.3480043
复制
发表时间:
2021-10
期刊:
MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Shiyu Li;Edward Hanson;Xuehai Qian;H. Li;Yiran Chen
Shiyu Li;Edward Hanson;Xuehai Qian;H. Li;Yiran Chen
中科院分区:
其他
文献类型:
--
作者:
Shiyu Li;Edward Hanson;Xuehai Qian;H. Li;Yiran Chen

文献摘要

相似文献

不断增长的参数大小和卷积神经网络(CNN)模型阻碍了他们的部署到资源受限的平台上,以消除CNN参数中的冗余,并提出稀疏的模型。那些效率设计人员通过硬件觉醒的算法来解决这个问题,但这些解决方案的修剪率在很大程度上限制了强制性的稀疏模式。并在不探索这些优势的情况下实现高压比,我们提出了基于算法分解的算法 - 硬件的共同设计方法。分解的重量。建筑水平,提议升级了一种新颖的“基础”数据流程及其相应的微体系结构设计,以最大程度地提高分解卷积带来的好处。 CIFAR-10和Imagenet上的模型的比率分别与以前的致密和稀疏加速器相比,升级的加速器升级可将能量效率提高8.3倍和3.77倍,并将潜伏期分别降低17.9××和2.16×。
The ever-growing parameter size and computation cost of Convolutional Neural Network (CNN) models hinder their deployment onto resource-constrained platforms. Network pruning techniques are proposed to remove the redundancy in CNN parameters and produce a sparse model. Sparse-aware accelerators are also proposed to reduce the computation cost and memory bandwidth requirements of inference by leveraging the model sparsity. The irregularity of sparse patterns, however, limits the efficiency of those designs. Researchers proposed to address this issue by creating a regular sparsity pattern through hardware-aware pruning algorithms. However, the pruning rate of these solutions is largely limited by the enforced sparsity patterns. This limitation motivates us to explore other compression methods beyond pruning. With two decoupled computation stages, we found that kernel decomposition could potentially take the processing of the sparse pattern off from the critical path of inference and achieve a high compression ratio without enforcing the sparse patterns. To exploit these advantages, we propose ESCALATE, an algorithm-hardware co-design approach based on kernel decomposition. At algorithm level, ESCALATE reorganizes the two computation stages of the decomposed convolution to enable a stream processing of the intermediate feature map. We proposed a hybrid quantization to exploit the different reuse frequency of each part of the decomposed weight. At architecture level, ESCALATE proposes a novel ‘Basis-First’ dataflow and its corresponding microarchitecture design to maximize the benefits brought by the decomposed convolution. We evaluate ESCALATE with four representative CNN models on both CIFAR-10 and ImageNet datasets and compare it against previous sparse accelerators and pruning algorithms. Results show that ESCALATE can achieve up to 325 × and 11 × compression ratio for models on CIFAR-10 and ImageNet, respectively. Comparing with previous dense and sparse accelerators, ESCALATE accelerator averagely boosts the energy efficiency by 8.3 × and 3.77 ×, and reduces the latency by 17.9 × and 2.16 ×, respectively.