DepthShrinker: A New Compression Paradigm Towards Boosting Real-Hardware Efficiency of Compact Neural Networks

DepthShrinker: A New Compression Paradigm Towards Boosting Real-Hardware Efficiency of Compact Neural Networks
复制标题

DOI:
10.48550/arxiv.2206.00843
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Y. Fu;Haichuan Yang;Jiayi Yuan;Meng Li;Cheng Wan;Raghuraman Krishnamoorthi;Vikas Chandra;Yingyan Lin
Y. Fu;Haichuan Yang;Jiayi Yuan;Meng Li;Cheng Wan;Raghuraman Krishnamoorthi;Vikas Chandra;Yingyan Lin
中科院分区:
其他
文献类型:
--
作者:
Y. Fu;Haichuan Yang;Jiayi Yuan;Meng Li;Cheng Wan;Raghuraman Krishnamoorthi;Vikas Chandra;Yingyan Lin

文献摘要

相似文献

配备紧致算子(例如,深度卷积)的高效深度神经网络(DNN)模型在保持良好的模型精度的同时,在降低DNN的理论复杂性(例如,权重/运算的总数)方面显示出巨大的潜力。然而,现有的高效DNN在实现其在提高真实硬件效率方面的承诺方面仍然有限,这是因为它们通常采用紧凑型运营商的低硬件利用率。在这项工作中,我们开辟了一种新的压缩范例来开发实硬件高效的DNN,从而在保持模型精度的同时提高了硬件效率。有趣的是,我们观察到,虽然一些DNN层的激活函数有助于DNN的训练优化和达到的精度,但它们可以在训练后适当地去除,而不会影响模型的精度。受此启发,我们提出了一个称为DepthShrinker的框架,该框架通过将现有高效DNN的基本构建块收缩成密集的计算模式来开发硬件友好的紧凑型网络,这些DNN具有不规则的计算模式,从而大大提高了硬件利用率和实际硬件效率。令人兴奋的是,我们的DepthShrinker框架提供了硬件友好的紧凑型网络,其性能优于最先进的高效DNN和压缩技术,例如,在Tesla V100上,准确率比SOTA通道修剪方法MetaPruning高3.06%,吞吐量是SOTA通道剪枝方法MetaPruning的1.53美元\倍。我们的代码请访问:https://github.com/facebookresearch/DepthShrinker.
Efficient deep neural network (DNN) models equipped with compact operators (e.g., depthwise convolutions) have shown great potential in reducing DNNs' theoretical complexity (e.g., the total number of weights/operations) while maintaining a decent model accuracy. However, existing efficient DNNs are still limited in fulfilling their promise in boosting real-hardware efficiency, due to their commonly adopted compact operators' low hardware utilization. In this work, we open up a new compression paradigm for developing real-hardware efficient DNNs, leading to boosted hardware efficiency while maintaining model accuracy. Interestingly, we observe that while some DNN layers' activation functions help DNNs' training optimization and achievable accuracy, they can be properly removed after training without compromising the model accuracy. Inspired by this observation, we propose a framework dubbed DepthShrinker, which develops hardware-friendly compact networks via shrinking the basic building blocks of existing efficient DNNs that feature irregular computation patterns into dense ones with much improved hardware utilization and thus real-hardware efficiency. Excitingly, our DepthShrinker framework delivers hardware-friendly compact networks that outperform both state-of-the-art efficient DNNs and compression techniques, e.g., a 3.06% higher accuracy and 1.53$\times$ throughput on Tesla V100 over SOTA channel-wise pruning method MetaPruning. Our codes are available at: https://github.com/facebookresearch/DepthShrinker.