Layer Freezing & Data Sieving: Missing Pieces of a Generic Framework for Sparse Training

Layer Freezing & Data Sieving: Missing Pieces of a Generic Framework for Sparse Training
复制标题

DOI:
10.48550/arxiv.2209.11204
复制
发表时间:
2022-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Geng Yuan;Yanyu Li;Sheng Li;Zhenglun Kong;S. Tulyakov;Xulong Tang;Yanzhi Wang;Jian Ren
Geng Yuan;Yanyu Li;Sheng Li;Zhenglun Kong;S. Tulyakov;Xulong Tang;Yanzhi Wang;Jian Ren
中科院分区:
其他
文献类型:
--
作者:
Geng Yuan;Yanyu Li;Sheng Li;Zhenglun Kong;S. Tulyakov;Xulong Tang;Yanzhi Wang;Jian Ren

文献摘要

相似文献

最近,稀疏训练已经成为边缘设备上高效深度学习的一个有前途的范例。目前的研究主要致力于通过进一步增加模型稀疏性来降低训练成本。然而,增加稀疏性并不总是理想的,因为它将不可避免地引入严重的精度退化在一个非常高的稀疏性水平。本文旨在探索其他可能的方向,以有效和高效地降低稀疏训练成本,同时保持准确性。为此,我们研究了两种技术,即层冻结和数据筛选。首先,层冻结方法在密集模型训练和微调方面取得了成功,但在稀疏训练领域从未采用过。然而,稀疏训练的独特特性可能会阻碍层冻结技术的结合。因此,我们分析了在稀疏训练中使用层冻结技术的可行性和潜力,发现它具有节省大量训练成本的潜力。第二,我们提出了一种数据筛选方法来进行有效的训练,通过确保在整个训练过程中只使用部分数据集,进一步降低了训练成本。我们表明,这两种技术可以很好地纳入稀疏训练算法,形成一个通用的框架,我们称之为SpFDE。我们广泛的实验表明,SpFDE可以显着降低训练成本,同时从三个维度保持准确性:权重稀疏性,层冻结和数据集筛选。
Recently, sparse training has emerged as a promising paradigm for efficient deep learning on edge devices. The current research mainly devotes efforts to reducing training costs by further increasing model sparsity. However, increasing sparsity is not always ideal since it will inevitably introduce severe accuracy degradation at an extremely high sparsity level. This paper intends to explore other possible directions to effectively and efficiently reduce sparse training costs while preserving accuracy. To this end, we investigate two techniques, namely, layer freezing and data sieving. First, the layer freezing approach has shown its success in dense model training and fine-tuning, yet it has never been adopted in the sparse training domain. Nevertheless, the unique characteristics of sparse training may hinder the incorporation of layer freezing techniques. Therefore, we analyze the feasibility and potentiality of using the layer freezing technique in sparse training and find it has the potential to save considerable training costs. Second, we propose a data sieving method for dataset-efficient training, which further reduces training costs by ensuring only a partial dataset is used throughout the entire training process. We show that both techniques can be well incorporated into the sparse training algorithm to form a generic framework, which we dub SpFDE. Our extensive experiments demonstrate that SpFDE can significantly reduce training costs while preserving accuracy from three dimensions: weight sparsity, layer freezing, and dataset sieving.