Structured Sparsity of Convolutional Neural Networks via Nonconvex Sparse Group Regularization

Structured Sparsity of Convolutional Neural Networks via Nonconvex Sparse Group Regularization
复制标题

DOI:
10.3389/fams.2020.529564
复制
发表时间:
2021-02
期刊:
--
影响因子:
--
通讯作者:
Kevin Bui;Fredrick Park;Shuai Zhang;Y. Qi;J. Xin
Kevin Bui;Fredrick Park;Shuai Zhang;Y. Qi;J. Xin
中科院分区:
其他
文献类型:
--
作者:
Kevin Bui;Fredrick Park;Shuai Zhang;Y. Qi;J. Xin

文献摘要

被引文献

相似文献

卷积神经网络(CNN)最近取得了巨大的成功,在各种成像应用中具有上级精度和性能,例如分类,对象检测和分割。然而,一个高度准确的CNN模型需要数百万个参数来训练和使用。即使稍微提高其性能,由于添加更多的层和/或增加每层的过滤器数量,也需要显著更多的参数。显然,许多权重参数都是冗余和无关的,因此原始的稠密模型可以被其压缩版本所取代,该压缩版本是通过在训练期间对层权重施加组间和组内稀疏性来获得的。在本文中,我们提出了一个非凸族的稀疏群套索,它混合了非凸正则化(例如,transformed(变换后的),它将稀疏性引入到各个权重上,并将正则化引入到层的输出通道上。我们将变量分裂应用到所提出的正则化中,以开发一种每次迭代由两个步骤组成的算法:梯度下降和阈值。在各种CNN架构上进行了数值实验,展示了稀疏群套索的非凸族在网络稀疏化和测试精度方面的有效性,与当前最先进的技术水平相当。
Convolutional neural networks (CNN) have been hugely successful recently with superior accuracy and performance in various imaging applications, such as classification, object detection, and segmentation. However, a highly accurate CNN model requires millions of parameters to be trained and utilized. Even to increase its performance slightly would require significantly more parameters due to adding more layers and/or increasing the number of filters per layer. Apparently, many of these weight parameters turn out to be redundant and extraneous, so the original, dense model can be replaced by its compressed version attained by imposing inter- and intra-group sparsity onto the layer weights during training. In this paper, we propose a nonconvex family of sparse group lasso that blends nonconvex regularization (e.g., transformed ℓ 1 , ℓ 1 − ℓ 2 , and ℓ 0 ) that induces sparsity onto the individual weights and ℓ 2,1 regularization onto the output channels of a layer. We apply variable splitting onto the proposed regularization to develop an algorithm that consists of two steps per iteration: gradient descent and thresholding. Numerical experiments are demonstrated on various CNN architectures showcasing the effectiveness of the nonconvex family of sparse group lasso in network sparsification and test accuracy on par with the current state of the art.