SmaQ: Smart Quantization for DNN Training by Exploiting Value Clustering
SmaQ: Smart Quantization for DNN Training by Exploiting Value Clustering
复制标题
DOI:
10.1109/lca.2021.3108505
复制
发表时间:
2021-07
影响因子:
2.3
通讯作者:
Nima Shoghi;Andrei Bersatti;Moinuddin K. Qureshi;Hyesoon Kim
中科院分区:
文献类型:
--
作者:
Nima Shoghi;Andrei Bersatti;Moinuddin K. Qureshi;Hyesoon Kim
Advancements in modern deep learning have shown that deeper networks with larger datasets can achieve state of the art results in many different tasks. As networks become deeper, the memory requirement of neural network training proves to be the primary bottleneck of single-machine training. In this letter, we first study the characteristics of neural network weight, gradient, feature map, gradient map, and optimizer state distributions for some popular neural network architectures. Our investigation shows that the majority of the data structures used by neural networks can have their value distributions be approximated with normal distributions. We then introduce Smart Quantization (SmaQ), a quantization scheme that exploits this observed normal distribution to quantize the data structures. Our dynamic quantization method calculates the sampled mean and standard deviation of tensors and quantizes each tensor element to 6 or 8 bits based on the z-score of that value. Our scheme reduces the memory usage during training by up to 6.7x with minor losses in accuracy.