Training Your Sparse Neural Network Better with Any Mask

Training Your Sparse Neural Network Better with Any Mask
复制标题

DOI:
10.48550/arxiv.2206.12755
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Ajay Jaiswal;Haoyu Ma;Tianlong Chen;Ying Ding;Zhangyang Wang
Ajay Jaiswal;Haoyu Ma;Tianlong Chen;Ying Ding;Zhangyang Wang
中科院分区:
其他
文献类型:
--
作者:
Ajay Jaiswal;Haoyu Ma;Tianlong Chen;Ying Ding;Zhangyang Wang

文献摘要

被引文献

相似文献

修剪大型神经网络以创建高质量的、可独立训练的稀疏掩码,其可以保持与其密集对应物相似的性能,由于降低了空间和时间复杂度,因此非常理想。由于研究工作的重点是越来越复杂的修剪方法,导致稀疏的子网络从零开始训练,我们主张一个正交的,探索不足的主题:改进修剪子网络的训练技术,即稀疏训练。除了普遍认为只有稀疏掩码的质量对稀疏训练很重要之外,在本文中,我们展示了一个替代的机会:人们可以仔细定制稀疏训练技术,以偏离默认的密集网络训练协议,包括在训练的早期阶段引入“幽灵”神经元和跳过连接,并策略性地修改初始化和标签。我们的新稀疏训练配方通常适用于使用各种稀疏掩码从头开始改进训练。通过采用我们新策划的技术,我们在各种流行的数据集(CIFAR-10,CIFAR-100,TinyImageNet),架构(ResNet-18/32/104,Vgg 16,MobileNet)和稀疏掩码选项(彩票,SNIP/GRASP,SynFlow,甚至随机修剪)上表现出显着的性能提升,与默认的训练协议相比,特别是在高稀疏度水平下。代码位于https://github.com/VITA-Group/ToST
Pruning large neural networks to create high-quality, independently trainable sparse masks, which can maintain similar performance to their dense counterparts, is very desirable due to the reduced space and time complexity. As research effort is focused on increasingly sophisticated pruning methods that leads to sparse subnetworks trainable from the scratch, we argue for an orthogonal, under-explored theme: improving training techniques for pruned sub-networks, i.e. sparse training. Apart from the popular belief that only the quality of sparse masks matters for sparse training, in this paper we demonstrate an alternative opportunity: one can carefully customize the sparse training techniques to deviate from the default dense network training protocols, consisting of introducing ``ghost"neurons and skip connections at the early stage of training, and strategically modifying the initialization as well as labels. Our new sparse training recipe is generally applicable to improving training from scratch with various sparse masks. By adopting our newly curated techniques, we demonstrate significant performance gains across various popular datasets (CIFAR-10, CIFAR-100, TinyImageNet), architectures (ResNet-18/32/104, Vgg16, MobileNet), and sparse mask options (lottery ticket, SNIP/GRASP, SynFlow, or even randomly pruning), compared to the default training protocols, especially at high sparsity levels. Code is at https://github.com/VITA-Group/ToST