HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks

HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
复制标题

DOI:
10.1609/aaai.v37i9.26244
复制
发表时间:
2023-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Jinqi Xiao;Chengming Zhang;Yu Gong;Miao Yin;Yang Sui;Lizhi Xiang;Dingwen Tao;Bo Yuan
Jinqi Xiao;Chengming Zhang;Yu Gong;Miao Yin;Yang Sui;Lizhi Xiang;Dingwen Tao;Bo Yuan
中科院分区:
其他
文献类型:
--
作者:
Jinqi Xiao;Chengming Zhang;Yu Gong;Miao Yin;Yang Sui;Lizhi Xiang;Dingwen Tao;Bo Yuan

文献摘要

相似文献

低秩压缩是获得紧凑神经网络模型的一种重要的模型压缩策略。一般来说,由于秩值直接决定模型的复杂性和模型的准确性,因此适当地选择分层秩是非常关键和期望的。到目前为止,虽然已经提出了许多低秩压缩方法,无论是手动或自动选择的行列,他们遭受昂贵的手动试验或不满意的压缩性能。此外,所有现有的作品都不是以硬件感知的方式设计的,限制了压缩模型在真实硬件平台上的实际性能。为了解决这些挑战,在本文中,我们提出了HALOC,一个硬件感知的自动低秩压缩框架。通过从架构搜索的角度解释自动排名选择,我们开发了一个端到端的解决方案,以确定合适的逐层排名在一个可区分的和硬件感知的方式。我们进一步提出设计原则和缓解策略,以有效地探索秩空间,减少潜在的干扰问题。在不同数据集和硬件平台上的实验结果表明了该方法的有效性。在CIFAR-10数据集上,HALOC比未压缩的ResNet-20和VGG-16模型分别提高了0.07%和0.38%的精度,FLOP分别减少了72.20%和86.44%。在ImageNet数据集上,HALOC的top-1准确率比原始ResNet-18模型高0.9%,FLOP减少66.16%。HALOC还显示出比最先进的自动低秩压缩解决方案高0.66%的top-1准确性增加,具有更少的计算和内存成本。此外,HALOC还展示了在不同硬件平台上的实际加速效果,并通过桌面GPU、嵌入式GPU和ASIC加速器上的测试结果进行了验证。
Low-rank compression is an important model compression strategy for obtaining compact neural network models. In general, because the rank values directly determine the model complexity and model accuracy, proper selection of layer-wise rank is very critical and desired. To date, though many low-rank compression approaches, either selecting the ranks in a manual or automatic way, have been proposed, they suffer from costly manual trials or unsatisfied compression performance. In addition, all of the existing works are not designed in a hardware-aware way, limiting the practical performance of the compressed models on real-world hardware platforms. To address these challenges, in this paper we propose HALOC, a hardware-aware automatic low-rank compression framework. By interpreting automatic rank selection from an architecture search perspective, we develop an end-to-end solution to determine the suitable layer-wise ranks in a differentiable and hardware-aware way. We further propose design principles and mitigation strategy to efficiently explore the rank space and reduce the potential interference problem. Experimental results on different datasets and hardware platforms demonstrate the effectiveness of our proposed approach. On CIFAR-10 dataset, HALOC enables 0.07% and 0.38% accuracy increase over the uncompressed ResNet-20 and VGG-16 models with 72.20% and 86.44% fewer FLOPs, respectively. On ImageNet dataset, HALOC achieves 0.9% higher top-1 accuracy than the original ResNet-18 model with 66.16% fewer FLOPs. HALOC also shows 0.66% higher top-1 accuracy increase than the state-of-the-art automatic low-rank compression solution with fewer computational and memory costs. In addition, HALOC demonstrates the practical speedups on different hardware platforms, verified by the measurement results on desktop GPU, embedded GPU and ASIC accelerator.