Model Preserving Compression for Neural Networks

Model Preserving Compression for Neural Networks
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
--
影响因子:
--
通讯作者:
Jerry Chee;Megan Flynn;Anil Damle;Chris De Sa
Jerry Chee;Megan Flynn;Anil Damle;Chris De Sa
中科院分区:
其他
文献类型:
--
作者:
Jerry Chee;Megan Flynn;Anil Damle;Chris De Sa

文献摘要

被引文献

相似文献

在训练复杂的深度学习模型后,一个常见的任务是压缩模型以减少计算和存储需求。在压缩时,我们希望保留原始模型的每例决策(例如,超越top-1精度或保持鲁棒性),保持网络的结构,自动确定每层压缩级别,并消除微调的需要。没有现有的压缩方法同时满足这些标准$\unicode{x2014}$我们引入了一种利用插值分解的原则方法。我们的方法同时选择和消除通道(类似地,神经元),然后构建一个插值矩阵,将校正传播到下一层,保持网络的结构。因此,我们的方法即使在没有微调的情况下也能获得良好的性能,并且可以进行理论分析。我们对单层网络的理论泛化界限自然地使我们的方法能够自动选择深度网络的每层大小。从简单的单隐藏层网络到ImageNet上的深度网络,我们在各种任务、模型和数据集$\unicode{x2014}$上展示了我们的方法的有效性。
After training complex deep learning models, a common task is to compress the model to reduce compute and storage demands. When compressing, it is desirable to preserve the original model's per-example decisions (e.g., to go beyond top-1 accuracy or preserve robustness), maintain the network's structure, automatically determine per-layer compression levels, and eliminate the need for fine tuning. No existing compression methods simultaneously satisfy these criteria $\unicode{x2014}$ we introduce a principled approach that does by leveraging interpolative decompositions. Our approach simultaneously selects and eliminates channels (analogously, neurons), then constructs an interpolation matrix that propagates a correction into the next layer, preserving the network's structure. Consequently, our method achieves good performance even without fine tuning and admits theoretical analysis. Our theoretical generalization bound for a one layer network lends itself naturally to a heuristic that allows our method to automatically choose per-layer sizes for deep networks. We demonstrate the efficacy of our approach with strong empirical performance on a variety of tasks, models, and datasets $\unicode{x2014}$ from simple one-hidden-layer networks to deep networks on ImageNet.