SmaQ: Smart Quantization for DNN Training by Exploiting Value Clustering

SmaQ: Smart Quantization for DNN Training by Exploiting Value Clustering
复制标题

DOI:
10.1109/lca.2021.3108505
复制
发表时间:
2021-07
影响因子:
2.3
通讯作者:
Nima Shoghi;Andrei Bersatti;Moinuddin K. Qureshi;Hyesoon Kim
Nima Shoghi;Andrei Bersatti;Moinuddin K. Qureshi;Hyesoon Kim
中科院分区:
计算机科学3区
文献类型:
--
作者:
Nima Shoghi;Andrei Bersatti;Moinuddin K. Qureshi;Hyesoon Kim

文献摘要

相似文献

现代深度学习的进步表明,具有更大数据集的更深网络可以在许多不同的任务中实现最先进的结果。随着网络变得越来越深,神经网络训练的内存需求被证明是单机训练的主要瓶颈。在这封信中,我们首先研究了一些流行的神经网络架构的神经网络权重,梯度,特征映射,梯度映射和优化器状态分布的特性。我们的研究表明,神经网络使用的大多数数据结构都可以用正态分布近似其值分布。然后,我们引入智能量化(SmaQ),利用这种观察到的正态分布来优化数据结构的量化方案。我们的动态量化方法计算张量的采样均值和标准差,并根据该值的z分数将每个张量元素量化为6或8位。我们的方案将训练过程中的内存使用量减少了6.7倍,准确性略有损失。
Advancements in modern deep learning have shown that deeper networks with larger datasets can achieve state of the art results in many different tasks. As networks become deeper, the memory requirement of neural network training proves to be the primary bottleneck of single-machine training. In this letter, we first study the characteristics of neural network weight, gradient, feature map, gradient map, and optimizer state distributions for some popular neural network architectures. Our investigation shows that the majority of the data structures used by neural networks can have their value distributions be approximated with normal distributions. We then introduce Smart Quantization (SmaQ), a quantization scheme that exploits this observed normal distribution to quantize the data structures. Our dynamic quantization method calculates the sampled mean and standard deviation of tensors and quantizes each tensor element to 6 or 8 bits based on the z-score of that value. Our scheme reduces the memory usage during training by up to 6.7x with minor losses in accuracy.