Activation Density based Mixed-Precision Quantization for Energy Efficient Neural Networks

Activation Density based Mixed-Precision Quantization for Energy Efficient Neural Networks
复制标题

基于激活密度的节能神经网络混合精度量化

DOI:
10.23919/date51398.2021.9474031
复制
发表时间:
2021
期刊:
2021 Design, Automation & Test in Europe Conference & Exhibition (DATE)
影响因子:
--
通讯作者:
P. Panda
P. Panda
中科院分区:
--
文献类型:
--
作者:
Karina Vasquez;Yeshwanth Venkatesha;Abhiroop Bhattacharjee;Abhishek Moitra;P. Panda

文献摘要

参考文献

被引文献

相似文献

随着神经网络在嵌入式设备中的广泛采用,越来越需要模型压缩技术来促进资源受限环境中的无缝部署。量化是产生最先进的模型压缩的常用方法之一。大多数量化方法采用经过充分训练的模型,然后应用不同的启发式方法来确定网络不同层的最佳位精度,最后重新训练网络以恢复准确度的任何下降。基于激活密度(层中非零激活的比例),我们提出了一种新的训练量化方法。我们的方法在训练过程中计算每层的最佳位宽/精度,从而产生具有竞争力精度的节能混合精度模型。由于我们在训练过程中逐步训练较低精度的模型,因此我们的方法以较低的训练复杂度产生最终的量化模型,并且还消除了重新训练的需要。我们在VGG 19/ResNet 18架构上对CIFAR-10,CIFAR-100,TinyImagenet等基准数据集进行了实验,并报告了准确性和能量估计。在我们的实验中,我们在估计的乘法和累加(MAC)减少方面实现了高达4.5倍的收益,同时将训练复杂度降低了50%。为了进一步评估我们所提出的方法的能源效益,我们开发了一个混合精度可扩展的内存处理(PIM)硬件加速器平台。硬件平台采用移位-相加功能,用于处理多位精度神经网络模型。在PIM平台上评估使用我们提出的方法获得的量化模型,与基线16位模型相比,能量减少了约5倍。此外,我们发现,与基线16位精度未修剪模型相比,将基于激活密度的量化与基于激活密度的修剪(均在训练期间进行)相结合,在PIM平台上分别为VGG 19和ResNet 18架构产生高达198倍和44倍的能量减少。
As neural networks gain widespread adoption in embedded devices, there is a growing need for model compression techniques to facilitate seamless deployment in resource-constrained environments. Quantization is one of the go-to methods yielding state-of-the-art model compression. Most quantization approaches take a fully trained model, then apply different heuristics to determine the optimal bit-precision for different layers of the network, and finally retrain the network to regain any drop in accuracy. Based on Activation Density-the proportion of non-zero activations in a layer-we propose a novel in-training quantization method. Our method calculates optimal bit-width/precision for each layer during training yielding an energy-efficient mixed precision model with competitive accuracy. Since we train lower precision models progressively during training, our approach yields the final quantized model at lower training complexity and also eliminates the need for re-training. We run experiments on benchmark datasets like CIFAR-10, CIFAR-100, TinyImagenet on VGG19/ResNet18 architectures and report the accuracy and energy estimates for the same. We achieve up to 4.5× benefit in terms of estimated multiply-and-accumulate (MAC) reduction while reducing the training complexity by 50% in our experiments. To further evaluate the energy benefits of our proposed method, we develop a mixed-precision scalable Process In Memory (PIM) hardware accelerator platform. The hardware platform incorporates shift-add functionality for handling multibit precision neural network models. Evaluating the quantized models obtained with our proposed method on the PIM platform yields about 5× energy reduction compared to baseline 16-bit models. Additionally, we find that integrating activation density based quantization with activation density based pruning (both conducted during training) yields up to ~ 198× and ~44× energy reductions for VGG19 and ResNet18 architectures respectively on PIM platform compared to baseline 16-bit precision, unpruned models.
Q-PIM:一种基于遗传算法的灵活 DNN 量化方法及其在内存处理平台中的应用
DOI: 10.1109/dac18072.2020.9218737
发表时间: 2020
期刊: Design Automation Conference
影响因子: --
作者:
Long, Yun;Lee, Edward;Kim, Daehyun;Mukhopadhyay, Saibal
通讯作者: Mukhopadhyay, Saibal