Pre-Defined Sparsity for Low-Complexity Convolutional Neural Networks

Pre-Defined Sparsity for Low-Complexity Convolutional Neural Networks
复制标题

DOI:
10.1109/tc.2020.2972520
复制
发表时间:
2020-01
影响因子:
3.7
通讯作者:
Souvik Kundu;M. Nazemi;M. Pedram;K. Chugg;P. Beerel
Souvik Kundu;M. Nazemi;M. Pedram;K. Chugg;P. Beerel
中科院分区:
计算机科学2区
文献类型:
--
作者:
Souvik Kundu;M. Nazemi;M. Pedram;K. Chugg;P. Beerel

文献摘要

相似文献

处理深度卷积神经网络的高能源成本阻碍了它们在嵌入式系统和物联网设备等能源受限平台中的普遍部署。本文介绍了具有预定义稀疏2D内核的卷积层,这些内核具有在过滤器内和过滤器之间周期性重复的支持集。由于我们的周期性稀疏内核的有效存储,由于减少了DRAM访问,参数节省可以转化为能源效率的显著提高,从而有望显着改善训练和推理的能耗和准确性之间的权衡。为了评估这种方法,我们在ResNet 18和VGG 16架构的稀疏变体中使用两个广泛接受的数据集CIFAR-10和Tiny ImageNet进行了实验。与基线模型相比,我们提出的稀疏变体需要的模型参数减少了82%,FLOP减少了5.6倍,ResNet 18在CIFAR-10上的准确性损失可以忽略不计。对于在Tiny ImageNet上训练的VGG 16,我们的方法需要减少5.8 × 5.8 × FLOPs,减少高达83.3%的模型参数,而前5名(前1名)的准确度仅下降1.2%(2.1%)。我们还比较了我们提出的架构与ShuffleNet和MobileNetV 2的性能。使用类似的超参数和FLOP,我们的ResNet 18变体产生了2.8%的平均准确性提高。
The high energy cost of processing deep convolutional neural networks impedes their ubiquitous deployment in energy-constrained platforms such as embedded systems and IoT devices. This article introduces convolutional layers with pre-defined sparse 2D kernels that have support sets that repeat periodically within and across filters. Due to the efficient storage of our periodic sparse kernels, the parameter savings can translate into considerable improvements in energy efficiency due to reduced DRAM accesses, thus promising significant improvements in the trade-off between energy consumption and accuracy for both training and inference. To evaluate this approach, we performed experiments with two widely accepted datasets, CIFAR-10 and Tiny ImageNet in sparse variants of the ResNet18 and VGG16 architectures. Compared to baseline models, our proposed sparse variants require up to $\mathord {\sim }82\%$∼82% fewer model parameters with $5.6\times$5.6× fewer FLOPs with negligible loss in accuracy for ResNet18 on CIFAR-10. For VGG16 trained on Tiny ImageNet, our approach requires $5.8 \times$5.8× fewer FLOPs and up to $\mathord {\sim }83.3\%$∼83.3% fewer model parameters with a drop in top-5 (top-1) accuracy of only 1.2% ($\mathord {\sim }2.1\%$∼2.1%). We also compared the performance of our proposed architectures with that of ShuffleNet and MobileNetV2. Using similar hyperparameters and FLOPs, our ResNet18 variants yield an average accuracy improvement of $\mathord {\sim }2.8\%$∼2.8%.