CSCNN: Algorithm-hardware Co-design for CNN Accelerators using Centrosymmetric Filters

CSCNN: Algorithm-hardware Co-design for CNN Accelerators using Centrosymmetric Filters
复制标题

DOI:
10.1109/hpca51647.2021.00058
复制
发表时间:
2021-02
期刊:
2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Jiajun Li;A. Louri;Avinash Karanth;Razvan C. Bunescu
Jiajun Li;A. Louri;Avinash Karanth;Razvan C. Bunescu
中科院分区:
其他
文献类型:
--
作者:
Jiajun Li;A. Louri;Avinash Karanth;Razvan C. Bunescu

文献摘要

相似文献

卷积神经网络(CNN)是计算机视觉、语音和文本处理中许多最先进的深度学习模型的核心。训练和部署这种基于CNN的架构通常需要大量的计算资源。稀疏性已经成为一种减少CNN数据量和计算量的有效压缩方法。然而,稀疏性往往会导致计算的不规则性,这阻碍了加速器充分利用其对性能和能量改进的好处。在本文中,我们提出了CSCNN,一个用于CNN压缩和加速的算法/硬件协同设计框架,它减轻了计算不规则性的影响,提供了更好的性能和能量效率。在算法方面,CSCNN使用中心对称矩阵作为卷积滤波器。通过这样做,它将所需权重的数量减少了近50%,并在不影响规则性和准确性的情况下实现了结构化计算重用。此外,利用互补的剪枝技术进一步减少了2.8-7.2倍的计算量,而精度损失微乎其微。在硬件方面,我们提出了一种CSCNN加速器,它有效地利用了中心对称过滤器所带来的结构化计算重用,并进一步消除了零计算,从而提高了性能和能源效率。与密集加速器SCNN和SparTen相比,该加速器的性能分别提高了3.7倍、1.6倍和1.3倍,能量延迟积分别提高了8.9倍、2.8倍和2.0倍。
Convolutional neural networks (CNNs) are at the core of many state-of-the-art deep learning models in computer vision, speech, and text processing. Training and deploying such CNN-based architectures usually require a significant amount of computational resources. Sparsity has emerged as an effective compression approach for reducing the amount of data and computation for CNNs. However, sparsity often results in computational irregularity, which prevents accelerators from fully taking advantage of its benefits for performance and energy improvement. In this paper, we propose CSCNN, an algorithm/hardware co-design framework for CNN compression and acceleration that mitigates the effects of computational irregularity and provides better performance and energy efficiency. On the algorithmic side, CSCNN uses centrosymmetric matrices as convolutional filters. In doing so, it reduces the number of required weights by nearly 50% and enables structured computational reuse without compromising regularity and accuracy. Additionally, complementary pruning techniques are leveraged to further reduce computation by a factor of $2.8-7.2\times $ with a marginal accuracy loss. On the hardware side, we propose a CSCNN accelerator that effectively exploits the structured computational reuse enabled by centrosymmetric filters, and further eliminates zero computations for increased performance and energy efficiency. Compared against a dense accelerator, SCNN and SparTen, the proposed accelerator performs $3.7\times $, $1.6\times $ and $1.3\times $ better, and improves the EDP (Energy Delay Product) by $8.9\times $, $2.8\times $ and $2.0\times $, respectively.