TDC: Towards Extremely Efficient CNNs on GPUs via Hardware-Aware Tucker Decomposition

TDC: Towards Extremely Efficient CNNs on GPUs via Hardware-Aware Tucker Decomposition
复制标题

DOI:
10.1145/3572848.3577478
复制
发表时间:
2022-11
期刊:
Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming
影响因子:
--
通讯作者:
Lizhi Xiang;Miao Yin;Chengming Zhang;Aravind Sukumaran-Rajam;P. Sadayappan;Bo Yuan;Dingwen Tao
Lizhi Xiang;Miao Yin;Chengming Zhang;Aravind Sukumaran-Rajam;P. Sadayappan;Bo Yuan;Dingwen Tao
中科院分区:
其他
文献类型:
--
作者:
Lizhi Xiang;Miao Yin;Chengming Zhang;Aravind Sukumaran-Rajam;P. Sadayappan;Bo Yuan;Dingwen Tao

文献摘要

相似文献

Tucker分解是SOTA CNN模型压缩技术之一,与减少的FLOP相比,我们使用现有的GPU软件(例如CUDNN)的tuckers型模型来缩短推理时间。基于ADMM的训练算法可以实现高度准确的Tucker-Format模型。代码在CUDNN上最多可达到2.21×速度,TVM上的1.12×速度以及使用Cudnn的原始型号上的3.27倍,最多最多为0.05%的精度损失。
Tucker decomposition is one of the SOTA CNN model compression techniques. However, unlike the FLOPs reduction, we observe very limited inference time reduction with Tucker-compressed models using existing GPU software such as cuDNN. To this end, we propose an efficient end-to-end framework that can generate highly accurate and compact CNN models via Tucker decomposition and optimized inference code on GPUs. Specifically, we propose an ADMM-based training algorithm that can achieve highly accurate Tucker-format models. We also develop a high-performance kernel for Tucker-format convolutions and analytical performance models to guide the selection of execution parameters. We further propose a co-design framework to determine the proper Tucker ranks driven by practical inference time (rather than FLOPs). Our evaluation on five modern CNNs with A100 demonstrates that our compressed models with our optimized code achieve up to 2.21× speedup over cuDNN, 1.12× speedup over TVM, and 3.27× over the original models using cuDNN with at most 0.05% accuracy loss.