Class-Discriminative CNN Compression

Class-Discriminative CNN Compression
复制标题

DOI:
10.1109/icpr56361.2022.9956066
复制
发表时间:
2021-10
期刊:
2022 26th International Conference on Pattern Recognition (ICPR)
影响因子:
--
通讯作者:
Yuchen Liu-;D. Wentzlaff;S. Kung
Yuchen Liu-;D. Wentzlaff;S. Kung
中科院分区:
其他
文献类型:
--
作者:
Yuchen Liu-;D. Wentzlaff;S. Kung

文献摘要

相似文献

通过修剪和蒸馏来压缩卷积神经网络(CNN)已经受到越来越多的关注。特别是,设计一种基于类别区分的方法是理想的,因为它与CNN的训练目标无缝契合。在本文中,我们提出了类判别压缩(CDC),它在修剪和蒸馏中注入了类判别,以促进CNN的训练目标。我们首先研究了一组判别函数对通道修剪的有效性,其中我们通过直观的概括在我们的研究中包括众所周知的单变量二进制类统计,如学生T检验。然后,我们提出了一种新的层自适应分层修剪方法,在那里我们使用一个粗糙的类歧视计划的早期层和一个罚款后的层。这种方法自然符合CNN在早期层处理粗糙语义并在后期提取精细概念的事实。此外,我们利用判别成分分析(DCA)提取知识的中间表示在子空间中具有丰富的判别信息,这提高了隐藏层的线性可分性和分类精度的学生。结合修剪和蒸馏,CDC在CIFAR和ILSVRC-2012上进行评估,我们始终优于最先进的结果。
Compressing convolutional neural networks (CNNs) by pruning and distillation has received ever-increasing focus. In particular, designing a class-discrimination based approach would be desired as it fits seamlessly with the CNNs training objective. In this paper, we propose class-discriminative compression (CDC), which injects class discrimination in both pruning and distillation to facilitate the CNNs training goal. We first study the effectiveness of a group of discriminant functions for channel pruning, where we include well-known single-variate binary-class statistics like Student’s T-Test in our study via an intuitive generalization. We then propose a novel layer-adaptive hierarchical pruning approach, where we use a coarse class discrimination scheme for early layers and a fine one for later layers. This method naturally accords with the fact that CNNs process coarse semantics in the early layers and extract fine concepts at the later. Moreover, we leverage discriminant component analysis (DCA) to distill knowledge of intermediate representations in a subspace with rich discriminative information, which enhances hidden layers’ linear separability and classification accuracy of the student. Combining pruning and distillation, CDC is evaluated on CIFAR and ILSVRC-2012, where we consistently outperform the state-of-the-art results.