An Efficient CNN Inference Accelerator Based on Intra- and Inter-Channel Feature Map Compression

An Efficient CNN Inference Accelerator Based on Intra- and Inter-Channel Feature Map Compression
复制标题

DOI:
10.1109/tcsi.2023.3287602
复制
发表时间:
2023-09
期刊:
IEEE Transactions on Circuits and Systems I: Regular Papers
影响因子:
--
通讯作者:
Chenjia Xie;Zhuang Shao;Ning Zhao;Yuan Du;Li Du
Chenjia Xie;Zhuang Shao;Ning Zhao;Yuan Du;Li Du
中科院分区:
其他
文献类型:
--
作者:
Chenjia Xie;Zhuang Shao;Ning Zhao;Yuan Du;Li Du

文献摘要

相似文献

深层卷积神经网络(CNN)在推理过程中产生密集的层间数据,这导致大量的片上存储空间和片外带宽。为了解决存储受限的问题,本文提出了一种采用压缩技术的加速器,通过去除通道内和通道间的冗余信息来减少层间数据。在压缩过程中使用主成分分析(PCA)来集中通道间信息。在每个特征图内部实施空间差异、截断和可重构位宽编码,以消除通道内数据冗余。此外,引入了一种特殊的数据排列来增强数据的连续性,从而优化了PCA分析,提高了压缩性能。采用所提出的压缩技术,设计了一个CNN加速器,通过流水线处理重建、CNN计算和压缩操作来支持即时压缩过程。该加速器样机采用28 nm CMOS工艺实现。在218.5 mW的功率下,达到了819.2GOPS的峰值吞吐量和3.75TOPS/W的能效。实验表明,该压缩技术对现有CNN的压缩比分别为21.5%、9.8%和19.3%,而准确率损失可以忽略不计。
Deep convolutional neural networks (CNNs) generate intensive inter-layer data during inference, which results in substantial on- chip memory size and off-chip bandwidth. To solve the memory constraint, this paper proposes an accelerator adopting a compression technique that can reduce the inter-layer data by removing both intra- and inter-channel redundant information. Principal component analysis (PCA) is utilized in the compression process to concentrate inter-channel information. The spatial differences, truncation, and reconfigurable bit-width coding are implemented inside every feature map to eliminate the intra-channel data redundancy. Moreover, a particular data arrangement is introduced to enhance data continuity to optimize PCA analysis and improve compression performance. A CNN accelerator with the proposed compression technique is designed to support the on- the-fly compression process by pipelining the reconstruction, CNN computation, and compression operation. The prototype accelerator is implemented using 28-nm CMOS technology. It achieves 819.2GOPS peak throughput and 3.75TOPS/W energy efficiency with 218.5mW. Experiments show that the proposed compression technique achieves compression ratios of 21.5% $\sim $ 43.0% (8-bit mode) and 9.8% $\sim $ 19.3% (16-bit mode) on state-of-the-art CNNs with a negligible accuracy loss.