Get More at Once: Alternating Sparse Training with Gradient Correction

Get More at Once: Alternating Sparse Training with Gradient Correction
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Li Yang;Jian Meng;J.-s. Seo;Deliang Fan
Li Yang;Jian Meng;J.-s. Seo;Deliang Fan
中科院分区:
其他
文献类型:
--
作者:
Li Yang;Jian Meng;J.-s. Seo;Deliang Fan

文献摘要

相似文献

最近,出现了探索训练稀疏性的新趋势,它在训练过程中删除参数,从而提高了训练和推理效率。这一系列的工作主要是为了获得一个单一的稀疏模型下预定义的大稀疏比。这导致静态/固定稀疏推理模型不能调整或重新配置其计算复杂度(即,推理结构、等待时间),以获得真实世界变化和动态的硬件资源可用性。为了实现这种运行时或训练后的网络变形,已经提出了“动态推理”或“一次性训练”的概念,以一次性训练由多个子网组成的单个网络,但每个子网可以以不同的计算复杂度执行相同的推理功能。然而,传统的动态推理训练方法需要一个多目标优化的联合训练方案,这遭受了非常大的训练开销。在这项工作中,我们第一次提出了一种新的交替稀疏训练(AST)方案,用于训练多个稀疏子网进行动态推理,与从头开始训练单个稀疏模型的情况相比,没有额外的训练成本。此外,为了在不损失优化泛化能力的前提下减小子网间权值更新的干扰,在组内迭代过程中提出梯度修正,以减小子网间权值更新的干扰.我们在多个数据集上对最先进的稀疏训练方法验证了所提出的AST,这表明AST达到了类似或更好的准确性,但只需要训练一次就可以得到具有不同稀疏率的多个稀疏子网。更重要的是,与传统的基于联合训练的动态推理训练方法相比,在不影响每个子网精度的情况下,完全消除了大量的训练开销。代码可在https://github.com/mengjian0502/AST上获得。
Recently, a new trend of exploring training sparsity has emerged, which removes parameters during training, leading to both training and inference efficiency improvement. This line of works primarily aims to obtain a single sparse model under a pre-defined large sparsity ratio. It leads to a static/fixed sparse inference model that is not capable of adjusting or re-configuring its computation complexity (i.e., inference structure, latency) after training for real-world varying and dynamic hardware resource availability. To enable such run-time or post-training network morphing, the concept of ‘dynamic inference’ or ‘training-once-for-all’ has been proposed to train a single network consisting of multiple sub-nets once, but each sub-net could perform the same inference function with different computing complexity. However, the traditional dynamic inference training method requires a joint training scheme with multi-objective optimization, which suffers from very large training overhead. In this work, for the first time, we propose a novel alternating sparse training (AST) scheme to train multiple sparse sub-nets for dynamic inference without extra training cost compared to the case of training a single sparse model from scratch. Furthermore, to mitigate the interference of weight update among sub-nets without losing the generalization of optimization, we pro-pose gradient correction within the inner-group iterations to reduce their weight update interference. We validate the proposed AST on multiple datasets against state-of-the-art sparse training methods, which shows that AST achieves similar or better accuracy, but only needs to train once to get multiple sparse sub-nets with different sparsity ratios. More importantly, comparing with the traditional joint training based dynamic inference training methodology, the large training overhead is completely eliminated without affecting the accuracy of each sub-net. Code is available at https://github.com/mengjian0502/AST .