Coarsening the Granularity: Towards Structurally Sparse Lottery Tickets

Coarsening the Granularity: Towards Structurally Sparse Lottery Tickets
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
--
影响因子:
--
通讯作者:
Tianlong Chen;Xuxi Chen;Xiaolong Ma;Yanzhi Wang;Zhangyang Wang
Tianlong Chen;Xuxi Chen;Xiaolong Ma;Yanzhi Wang;Zhangyang Wang
中科院分区:
其他
文献类型:
--
作者:
Tianlong Chen;Xuxi Chen;Xiaolong Ma;Yanzhi Wang;Zhangyang Wang

文献摘要

相似文献

彩票假设(LTH)表明,密集模型包含高度稀疏的子网络(即中奖彩票),可以单独训练以达到完全精度。尽管做出了许多令人兴奋的努力,但有一个“常识”很少受到挑战:中奖彩票是通过迭代幅度修剪(IMP)找到的,因此得到的修剪子网只有非结构化的稀疏性。这个差距限制了实际中奖的吸引力,因为高度不规则的稀疏模式很难在硬件上加速。同时,在IMP中直接用结构化修剪代替非结构化修剪对性能的损害更大,通常无法找到中奖票。在本文中,我们证明了在一般情况下可以有效地找到结构稀疏中奖票的第一个积极结果。其核心思想是在每一轮(非结构化)IMP之后附加“后处理技术”,以加强结构稀疏性的形成。具体地说,我们首先在一些被认为是重要的通道中“重新填充”被修剪过的元素,然后“重新分组”非零元素,以创建灵活的分组明智的结构模式。我们确定的通道和组明智的结构子网都赢得了彩票,现有硬件很容易支持实质性的推理速度。在跨多个网络骨干网的不同数据集上进行的大量实验一致地验证了我们的建议,表明LTH的硬件加速障碍现已被消除。具体来说,结构中奖票在{CIFAR, Tiny-ImageNet, ImageNet}上以{36%~80%,74%,58%}的稀疏度节省了高达{64.93%,64.84%,60.23%}的运行时间,同时保持了相当的准确性。代码在https://github.com/VITA-Group/Structure-LTH。
The lottery ticket hypothesis (LTH) has shown that dense models contain highly sparse subnetworks (i.e., winning tickets) that can be trained in isolation to match full accuracy. Despite many exciting efforts being made, there is one"commonsense"rarely challenged: a winning ticket is found by iterative magnitude pruning (IMP) and hence the resultant pruned subnetworks have only unstructured sparsity. That gap limits the appeal of winning tickets in practice, since the highly irregular sparse patterns are challenging to accelerate on hardware. Meanwhile, directly substituting structured pruning for unstructured pruning in IMP damages performance more severely and is usually unable to locate winning tickets. In this paper, we demonstrate the first positive result that a structurally sparse winning ticket can be effectively found in general. The core idea is to append"post-processing techniques"after each round of (unstructured) IMP, to enforce the formation of structural sparsity. Specifically, we first"re-fill"pruned elements back in some channels deemed to be important, and then"re-group"non-zero elements to create flexible group-wise structural patterns. Both our identified channel- and group-wise structural subnetworks win the lottery, with substantial inference speedups readily supported by existing hardware. Extensive experiments, conducted on diverse datasets across multiple network backbones, consistently validate our proposal, showing that the hardware acceleration roadblock of LTH is now removed. Specifically, the structural winning tickets obtain up to {64.93%, 64.84%, 60.23%} running time savings at {36%~80%, 74%, 58%} sparsity on {CIFAR, Tiny-ImageNet, ImageNet}, while maintaining comparable accuracy. Code is at https://github.com/VITA-Group/Structure-LTH.