Rate Distortion For Model Compression: From Theory To Practice

Rate Distortion For Model Compression: From Theory To Practice
复制标题

DOI:
--
复制
发表时间:
2018-10
期刊:
--
影响因子:
--
通讯作者:
Weihao Gao;Chong Wang;Sewoong Oh
Weihao Gao;Chong Wang;Sewoong Oh
中科院分区:
其他
文献类型:
--
作者:
Weihao Gao;Chong Wang;Sewoong Oh

文献摘要

相似文献

现代深度神经网络的巨大规模使得在内存和通信有限的情况下部署这些模型具有挑战性。因此,在不显著降低性能的情况下压缩训练好的模型已成为一项越来越重要的任务。最近取得了巨大的进步,其中主要的技术构建模块是参数修剪,参数共享(量化)和低秩分解。在本文中,我们提出了原则性的方法来改进这些构建块中使用的常见启发式方法,即修剪和量化。我们首先通过率失真理论研究了模型压缩的基本极限。我们将速率失真函数从数据压缩引入到模型压缩中来量化这一基本限制。我们证明了速率失真函数的下界,并证明了它在线性模型上的可实现性。尽管这种可实现的压缩方案在实践中难以实现,但本文的分析激发了一种新的模型压缩框架。该框架为模型压缩提供了一个新的目标函数,可以与其他类型的模型压缩器(如剪枝或量化)一起应用。从理论上证明了该方案对于压缩单隐层ReLU神经网络是最优的。经验表明,我们提出的方案在压缩精度权衡的基线上有所改进。
The enormous size of modern deep neural networks makes it challenging to deploy those models in memory and communication limited scenarios. Thus, compressing a trained model without a significant loss in performance has become an increasingly important task. Tremendous advances has been made recently, where the main technical building blocks are parameter pruning, parameter sharing (quantization), and low-rank factorization. In this paper, we propose principled approaches to improve upon the common heuristics used in those building blocks, namely pruning and quantization. We first study the fundamental limit for model compression via the rate distortion theory. We bring the rate distortion function from data compression to model compression to quantify this fundamental limit. We prove a lower bound for the rate distortion function and prove its achievability for linear models. Although this achievable compression scheme is intractable in practice, this analysis motivates a novel model compression framework. This framework provides a new objective function in model compression, which can be applied together with other classes of model compressor such as pruning or quantization. Theoretically, we prove that the proposed scheme is optimal for compressing one-hidden-layer ReLU neural networks. Empirically, we show that the proposed scheme improves upon the baseline in the compression-accuracy tradeoff.