BATUDE: Budget-Aware Neural Network Compression Based on Tucker Decomposition

BATUDE: Budget-Aware Neural Network Compression Based on Tucker Decomposition
复制标题

DOI:
10.1609/aaai.v36i8.20869
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Miao Yin;Huy Phan;Xiao Zang;Siyu Liao;Bo Yuan
Miao Yin;Huy Phan;Xiao Zang;Siyu Liao;Bo Yuan
中科院分区:
其他
文献类型:
--
作者:
Miao Yin;Huy Phan;Xiao Zang;Siyu Liao;Bo Yuan

文献摘要

被引文献

相似文献

模型压缩对于在资源受限的设备上高效部署深度神经网络(DNN)模型非常重要。在各种模型压缩方法中,高阶张量分解是特别有吸引力和有用的,因为分解后的模型非常小且结构完整。对于这类方法,张量秩是最重要的超参数,直接决定了压缩DNN模型的架构和任务性能。然而,作为一个NP-难问题,选择最佳张量排名下所需的预算是非常具有挑战性的,最先进的研究遭受不满意的压缩性能和耗时的搜索过程。为了系统地解决这个基本问题,在本文中,我们提出了BATUDE,这是一种基于预算感知的TUcker DEcomposition压缩方法,可以通过一次性训练有效地计算最佳张量秩。通过将秩选择过程集成到具有指定压缩预算的DNN训练过程中,从数据中学习DNN模型的张量秩,从而为压缩模型的压缩比和分类精度带来非常显着的改善。在ImageNet数据集上的实验结果表明,与未压缩的ResNet-18模型相比,我们的方法具有0.33%的前5高精度,计算成本降低2.52倍。对于ResNet-50,与未压缩模型相比,所提出的方法分别使前5名的准确度提高了0.37%和0.55%,计算成本降低了2.97倍和2.04倍。
Model compression is very important for the efficient deployment of deep neural network (DNN) models on resource-constrained devices. Among various model compression approaches, high-order tensor decomposition is particularly attractive and useful because the decomposed model is very small and fully structured. For this category of approaches, tensor ranks are the most important hyper-parameters that directly determine the architecture and task performance of the compressed DNN models. However, as an NP-hard problem, selecting optimal tensor ranks under the desired budget is very challenging and the state-of-the-art studies suffer from unsatisfied compression performance and timing-consuming search procedures. To systematically address this fundamental problem, in this paper we propose BATUDE, a Budget-Aware TUcker DEcomposition-based compression approach that can efficiently calculate optimal tensor ranks via one-shot training. By integrating the rank selecting procedure to the DNN training process with a specified compression budget, the tensor ranks of the DNN models are learned from the data and thereby bringing very significant improvement on both compression ratio and classification accuracy for the compressed models. The experimental results on ImageNet dataset show that our method enjoys 0.33% top-5 higher accuracy with 2.52X less computational cost as compared to the uncompressed ResNet-18 model. For ResNet-50, the proposed approach enables 0.37% and 0.55% top-5 accuracy increase with 2.97X and 2.04X computational cost reduction, respectively, over the uncompressed model.