Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding

Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding
复制标题

DOI:
--
复制
发表时间:
2015-10
期刊:
arXiv: Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Song Han;Huizi Mao;W. Dally
Song Han;Huizi Mao;W. Dally
中科院分区:
其他
文献类型:
--
作者:
Song Han;Huizi Mao;W. Dally

文献摘要

被引文献

相似文献

神经网络是计算密集型和内存密集型的,这使得它们难以部署在硬件资源有限的嵌入式系统上。为了解决这一限制,我们引入了“深度压缩”,这是一个三级管道:修剪,训练量化和霍夫曼编码,它们共同工作,将神经网络的存储需求减少了35倍至49倍,而不影响其准确性。我们的方法首先通过只学习重要的连接来修剪网络。接着,我们将权值合并以加强权值分享,最后,我们应用霍夫曼编码。在前两步之后,我们重新训练网络,以微调剩余的连接和量化的质心。修剪将连接的数量减少了9倍到13倍;量化则将表示每个连接的位数从32减少到5。在ImageNet数据集上,我们的方法将AlexNet所需的存储空间减少了35倍,从240 MB减少到6.9MB,而不损失准确性。我们的方法将VGG-16的大小从552 MB减少到11.3MB,并且没有损失准确性。这允许将模型适配到片上SRAM缓存中,而不是片外DRAM存储器中。我们的压缩方法也有利于在应用程序的大小和下载带宽受限的移动的应用程序中使用复杂的神经网络。以CPU、GPU和移动的GPU为基准,压缩网络具有3倍至4倍的分层加速和3倍至7倍的能效。
Neural networks are both computationally intensive and memory intensive, making them difficult to deploy on embedded systems with limited hardware resources. To address this limitation, we introduce "deep compression", a three stage pipeline: pruning, trained quantization and Huffman coding, that work together to reduce the storage requirement of neural networks by 35x to 49x without affecting their accuracy. Our method first prunes the network by learning only the important connections. Next, we quantize the weights to enforce weight sharing, finally, we apply Huffman coding. After the first two steps we retrain the network to fine tune the remaining connections and the quantized centroids. Pruning, reduces the number of connections by 9x to 13x; Quantization then reduces the number of bits that represent each connection from 32 to 5. On the ImageNet dataset, our method reduced the storage required by AlexNet by 35x, from 240MB to 6.9MB, without loss of accuracy. Our method reduced the size of VGG-16 by 49x from 552MB to 11.3MB, again with no loss of accuracy. This allows fitting the model into on-chip SRAM cache rather than off-chip DRAM memory. Our compression method also facilitates the use of complex neural networks in mobile applications where application size and download bandwidth are constrained. Benchmarked on CPU, GPU and mobile GPU, compressed network has 3x to 4x layerwise speedup and 3x to 7x better energy efficiency.