An Image Enhancing Pattern-based Sparsity for Real-time Inference on Mobile Devices

An Image Enhancing Pattern-based Sparsity for Real-time Inference on Mobile Devices
复制标题

DOI:
10.1007/978-3-030-58601-0_37
复制
发表时间:
2020-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Xiaolong Ma;Wei Niu;Tianyun Zhang;Sijia Liu;Fu-Ming Guo;Sheng Lin;Hongjia Li;Xiang Chen;Jian Tang;Kaisheng Ma;Bin Ren;Yanzhi Wang
Xiaolong Ma;Wei Niu;Tianyun Zhang;Sijia Liu;Fu-Ming Guo;Sheng Lin;Hongjia Li;Xiang Chen;Jian Tang;Kaisheng Ma;Bin Ren;Yanzhi Wang
中科院分区:
其他
文献类型:
--
作者:
Xiaolong Ma;Wei Niu;Tianyun Zhang;Sijia Liu;Fu-Ming Guo;Sheng Lin;Hongjia Li;Xiang Chen;Jian Tang;Kaisheng Ma;Bin Ren;Yanzhi Wang

文献摘要

相似文献

权值修剪是消除深度神经网络冗余的一种简单有效的方法,可以在各种平台上实现加速。然而,大多数修剪技术本质上是在模型精度和规则性之间进行权衡,这导致推理精度受损和限制设备上的加速性能。为了解决这个问题,我们引入了一个新的稀疏性维度,即基于模式的稀疏性,它包括模式稀疏性和连接稀疏性,并且变得高度精确和硬件友好。通过对模式的精心设计,本文提出的剪枝框架在不同的DNN结构和数据集上实现了前所未有的一致性和准确性的提高和更好的特征提取能力,并且我们的模式感知剪枝框架还同时实现了模式库提取、模式选择、模式和连通性剪枝和权值训练。我们在新的基于模式的稀疏性上的方法自然适合于在移动平台上高效执行DNN的编译器优化。据我们所知,这是移动设备第一次实现大规模深度神经网络模型的实时推理,这得益于基于模式的稀疏性的独特空间特性和编译器的代码生成能力。
Weight pruning has been widely acknowledged as a straightforward and effective method to eliminate redundancy in Deep Neural Networks (DNN), thereby achieving acceleration on various platforms. However, most of the pruning techniques are essentially trade-offs between model accuracy and regularity which lead to impaired inference accuracy and limited on-device acceleration performance. To solve the problem, we introduce a new sparsity dimension, namely pattern-based sparsity that comprises pattern and connectivity sparsity, and becoming both highly accurate and hardware friendly. With carefully designed patterns, the proposed pruning unprecedentedly and consistently achieves accuracy enhancement and better feature extraction ability on different DNN structures and datasets, and our pattern-aware pruning framework also achieves pattern library extraction, pattern selection, pattern and connectivity pruning and weight training simultaneously. Our approach on the new pattern-based sparsity naturally fits into compiler optimization for highly efficient DNN execution on mobile platforms. To the best of our knowledge, it is the first time that mobile devices achieve real-time inference for the large-scale DNN models thanks to the unique spatial property of pattern-based sparsity and the help of the code generation capability of compilers.