T1000: Mitigating the memory footprint of convolution neural networks with decomposition and re-fusion

T1000: Mitigating the memory footprint of convolution neural networks with decomposition and re-fusion
复制标题

T1000:通过分解和重新融合减少卷积神经网络的内存占用

DOI:
10.1016/j.future.2018.02.024
复制
发表时间:
2018-07
期刊:
Future Generation Computer Systems
影响因子:
--
通讯作者:
Depei Qian
Depei Qian
中科院分区:
其他
文献类型:
--
作者:
Changxi Liu;Hailong Yang;Rui Wang;Zhongzhi Luan;Depei Qian

文献摘要

参考文献

相似文献

近年来,卷积神经网络因其具有良好的准确性而显著地推进了计算机视觉和其他智能应用的前沿。然而,卷积层越深,精度越高,计算复杂度越高,这阻碍了其在嵌入式和移动等资源受限系统中的应用。虽然已有研究通过张量分解来降低卷积神经网络的计算复杂度,但由于张量分解产生的中间数据量急剧增长,消耗了大量的内存资源,现有工作尚未解决这一问题。在这项工作中,我们提出T1000在将规范多进分解应用于常规卷积层之后重新融合张量之间的卷积,这样我们可以获得降低计算复杂度的好处,同时减轻中间数据的内存占用。我们通过对两个著名的卷积神经网络AlexNet和VGG-19的卷积层应用正则多进分解和再融合来证明我们方法的有效性。与默认的规范化多进分解相比,我们的方法将AlexNet和VGG-19的中间数据的内存占用分别减少了84.6%和77.4%。此外,我们的方法将AlexNet和VGG-19的性能分别提高了1.77倍和1.4倍。
In recent years, convolution neural networks have significantly advanced the frontier of computer vision and other intelligent applications due to its promising accuracy. However, the improved accuracy comes with the formidable computation complexity with deeper convolution layers, which prevents its adoption on resource constrained system such as embedded and mobile. Although research efforts have been devoted to reduce the computation complexity of convolution neural networks through tensor decomposition, the volume of intermediate data generated by the tensor decomposition grows dramatically, which consumes more memory resource and has not been addressed by existing work. In this work, we propose T1000 to re-fuse the convolutions across tensors after applying the canonical polyadic decomposition to conventional convolution layers so that we can receive the benefit of reduced computation complexity, in the meanwhile mitigate the memory occupancy of the intermediate data. We demonstrate the effectiveness of our approach by applying canonical polyadic decomposition and re-fusion to the convolution layers of two well-known convolution neural networks, AlexNet and VGG-19 implemented with Caffe. Compared to the default canonical polyadic decomposition, our approach reduces the memory occupancy of the intermediate data by 84.6% and 77.4% for AlexNet and VGG-19 respectively. In addition, our approach improves the performance of AlexNet and VGG-19 by 1.77× and 1.4× respectively.
DOI: 10.1016/j.jnca.2017.09.001
发表时间: 2018-02
期刊: J. Netw. Comput. Appl.
影响因子: --
作者:
Weishan Zhang;Gaowa Wulan;Jia Zhai;Liang Xu;Dehai Zhao;Xin Liu;Su Yang;Jiehan Zhou
通讯作者: Weishan Zhang;Gaowa Wulan;Jia Zhai;Liang Xu;Dehai Zhao;Xin Liu;Su Yang;Jiehan Zhou
DOI: --
发表时间: 2013-12
期刊: --
影响因子: --
作者:
Jimmy Ba;R. Caruana
通讯作者: Jimmy Ba;R. Caruana
DOI: 10.1109/tmm.2016.2642789
发表时间: 2017-05-01
影响因子: 7.3
作者:
Li, Jianan;Wei, Yunchao;Yan, Shuicheng
通讯作者: Yan, Shuicheng
DOI: 10.1109/micro.2016.7783725
发表时间: 2016-10
期刊: 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子: --
作者:
Manoj Alwani;Han Chen;M. Ferdman;Peter Milder
通讯作者: Manoj Alwani;Han Chen;M. Ferdman;Peter Milder
DOI: 10.1007/978-3-319-22186-1_65
发表时间: 2015-08
期刊: --
影响因子: --
作者:
D. Al-Jumeily;A. Hussain;P. Fergus;Naeem Radi
通讯作者: D. Al-Jumeily;A. Hussain;P. Fergus;Naeem Radi