ASLog: An Area-Efficient CNN Accelerator for Per-Channel Logarithmic Post-Training Quantization

ASLog: An Area-Efficient CNN Accelerator for Per-Channel Logarithmic Post-Training Quantization
复制标题

DOI:
10.1109/tcsi.2023.3315299
复制
发表时间:
2023-12
期刊:
IEEE Transactions on Circuits and Systems I: Regular Papers
影响因子:
--
通讯作者:
Jiawei Xu;Jiangshan Fan;Baolin Nan;Chen Ding;Li-Rong Zheng;Z. Zou;Y. Huan
Jiawei Xu;Jiangshan Fan;Baolin Nan;Chen Ding;Li-Rong Zheng;Z. Zou;Y. Huan
中科院分区:
其他
文献类型:
--
作者:
Jiawei Xu;Jiangshan Fan;Baolin Nan;Chen Ding;Li-Rong Zheng;Z. Zou;Y. Huan

文献摘要

相似文献

训练后量化(PTQ)已被证明是卷积神经网络(CNN)的有效模型压缩技术,无需重新训练或访问标记数据集。然而,CNN加速器要实现PTQ方法的效率潜力仍然具有挑战性。大量的PTQ技术盲目追求理论上的高压缩效果和精度,而忽略了它们对实际硬件实现的影响,造成硬件开销大于效益。本文介绍了ASLog,一个PTQ友好的CNN加速器,以算法-硬件协同优化的方式探索了四个关键设计:第一个实用的4位对数PTQ流水线SLogII,无乘法器算术元件(AE)设计,节能的偏置校正元件(BCE)设计,以及每通道量化友好(PCF)架构和低功耗。所提出的SLogII PTQ流水线可以将对数PTQ的极限推到4位,与普通的8位乘法器相比,功耗和面积消耗降低了40%。本文提出的BCE和PCF设计是第一个考虑广泛使用的每通道量化和偏置校正技术的硬件影响,实现了一个高效的PTQ友好的实现与小的硬件开销。ASLog采用UMC 40 nm工艺验证,能效为12.2 TOPS/W,核心面积为0.80 mm 2。ASLog可以实现336.3 GOPS/mm 2的面积效率和>500 OPs/Byte的操作强度,与以前的相关工作相比,分别提高了1.85倍和1.12倍。
Post-training quantization (PTQ) has been proven an efficient model compression technique for Convolution Neural Networks (CNNs), without re-training or access to labeled datasets. However, it remains challenging for a CNN accelerator to fulfill the efficiency potential of PTQ methods. A large number of PTQ techniques blindly pursue high theoretic compression effect and accuracy, ignoring their impact on the actual hardware implementation, which causes more hardware overhead than benefit. This paper introduces ASLog, a PTQ-friendly CNN accelerator that explores four key designs in an algorithm-hardware co-optimizing manner: the first practical 4-bit logarithmic PTQ pipeline SLogII, the multiplier-free arithmetic element (AE) design, the energy-efficient bias correction element (BCE) design, and the per-channel quantization friendly (PCF) architecture and dataflow. The proposed SLogII PTQ pipeline can push the limit of logarithmic PTQ to 4-bit with 40% lower in power and area consumption compared with a common 8-bit multiplier. The BCE and PCF design proposed in this paper are the first to consider the hardware impact of the widely-used per-channel quantization and bias correction technique, enabling an efficient PTQ-friendly implementation with a small hardware overhead. The ASLog is validated in a UMC 40-nm process, with 12.2 TOPS/W energy efficiency and 0.80 mm2 core area. The ASLog can achieve 336.3 GOPS/mm2 area efficiency and >500 OPs/Byte operational intensity, which map to over $1.85\times $ and $1.12\times $ improvement compared with the previous related works.