Minimal Random Code Learning: Getting Bits Back from Compressed Model Parameters

Minimal Random Code Learning: Getting Bits Back from Compressed Model Parameters
复制标题

DOI:
--
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Marton Havasi;Robert Peharz;José Miguel Hernández-Lobato
Marton Havasi;Robert Peharz;José Miguel Hernández-Lobato
中科院分区:
其他
文献类型:
--
作者:
Marton Havasi;Robert Peharz;José Miguel Hernández-Lobato

文献摘要

被引文献

相似文献

虽然深度神经网络是一个非常成功的模型类,但它们的大内存占用对能耗、通信带宽和存储要求造成了相当大的压力。因此,减少模型大小已成为深度学习的终极目标。一种典型的方法是训练一组确定性权重,同时应用某些技术,如修剪和量化,以便经验权重分布变得符合香农式编码方案。然而,如本文所示,放松权重确定性并使用权重的完全变分分布可以实现更有效的编码方案,从而获得更高的压缩率。特别是,在经典的位回参数之后,我们使用随机样本对网络权重进行编码,仅需要与采样变分分布和编码分布之间的Kullback-Leibler散度对应的位数。通过对Kullback-Leibler散度施加约束,我们能够显式地控制压缩率,同时优化训练集上的预期损失。所采用的编码方案可以被证明是接近最佳的信息理论的下限,相对于所采用的变分家庭。我们的方法在神经网络压缩方面开创了新的最先进技术,因为它在帕累托意义上严格优于以前的方法:在基准LeNet-5/MNIST和VGG-16/CIFAR-10上,我们的方法在固定内存预算下产生了最佳测试性能,反之亦然,它在固定测试性能下实现了最高的压缩率。
While deep neural networks are a highly successful model class, their large memory footprint puts considerable strain on energy consumption, communication bandwidth, and storage requirements. Consequently, model size reduction has become an utmost goal in deep learning. A typical approach is to train a set of deterministic weights, while applying certain techniques such as pruning and quantization, in order that the empirical weight distribution becomes amenable to Shannon-style coding schemes. However, as shown in this paper, relaxing weight determinism and using a full variational distribution over weights allows for more efficient coding schemes and consequently higher compression rates. In particular, following the classical bits-back argument, we encode the network weights using a random sample, requiring only a number of bits corresponding to the Kullback-Leibler divergence between the sampled variational distribution and the encoding distribution. By imposing a constraint on the Kullback-Leibler divergence, we are able to explicitly control the compression rate, while optimizing the expected loss on the training set. The employed encoding scheme can be shown to be close to the optimal information-theoretical lower bound, with respect to the employed variational family. Our method sets new state-of-the-art in neural network compression, as it strictly dominates previous approaches in a Pareto sense: On the benchmarks LeNet-5/MNIST and VGG-16/CIFAR-10, our approach yields the best test performance for a fixed memory budget, and vice versa, it achieves the highest compression rates for a fixed test performance.