SparCL: Sparse Continual Learning on the Edge

SparCL: Sparse Continual Learning on the Edge
复制标题

DOI:
10.48550/arxiv.2209.09476
复制
发表时间:
2022-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Zifeng Wang;Zheng Zhan;Yifan Gong;Geng Yuan;Wei Niu;T. Jian;Bin Ren;Stratis Ioannidis;Yanzhi Wang;Jennifer G. Dy
Zifeng Wang;Zheng Zhan;Yifan Gong;Geng Yuan;Wei Niu;T. Jian;Bin Ren;Stratis Ioannidis;Yanzhi Wang;Jennifer G. Dy
中科院分区:
其他
文献类型:
--
作者:
Zifeng Wang;Zheng Zhan;Yifan Gong;Geng Yuan;Wei Niu;T. Jian;Bin Ren;Stratis Ioannidis;Yanzhi Wang;Jennifer G. Dy

文献摘要

相似文献

持续学习(CL)方面的现有工作侧重于减轻灾难性遗忘,即在学习新任务时,过去任务的模型性能下降。然而,CL 系统的训练效率尚未得到充分研究,这限制了 CL 系统在资源有限的场景下的实际应用。在这项工作中,我们提出了一种名为稀疏持续学习(SparCL)的新颖框架,这是第一个利用稀疏性在边缘设备上实现经济有效的持续学习的研究。 SparCL通过权重稀疏性、数据效率和梯度稀疏性三个方面的协同,实现了训练加速和精度保持。具体来说,我们提出任务感知动态掩蔽(TDM)来在整个 CL 过程中学习稀疏网络,动态数据删除(DDR)来删除信息量较少的训练数据,以及动态梯度掩蔽(DGM)来稀疏梯度更新。它们中的每一个不仅提高了效率,而且还进一步减少了灾难性遗忘。 SparCL 持续提高了现有最先进 (SOTA) CL 方法的训练效率,最多将训练 FLOP 减少 23 倍,并且令人惊讶的是,进一步将 SOTA 准确率提高最多 1.7%。 SparCL 在效率和准确性方面也优于通过将 SOTA 稀疏训练方法适应 CL 设置而获得的竞争基线。我们还评估了 SparCL 在真实手机上的有效性,进一步表明了我们方法的实际潜力。
Existing work in continual learning (CL) focuses on mitigating catastrophic forgetting, i.e., model performance deterioration on past tasks when learning a new task. However, the training efficiency of a CL system is under-investigated, which limits the real-world application of CL systems under resource-limited scenarios. In this work, we propose a novel framework called Sparse Continual Learning(SparCL), which is the first study that leverages sparsity to enable cost-effective continual learning on edge devices. SparCL achieves both training acceleration and accuracy preservation through the synergy of three aspects: weight sparsity, data efficiency, and gradient sparsity. Specifically, we propose task-aware dynamic masking (TDM) to learn a sparse network throughout the entire CL process, dynamic data removal (DDR) to remove less informative training data, and dynamic gradient masking (DGM) to sparsify the gradient updates. Each of them not only improves efficiency, but also further mitigates catastrophic forgetting. SparCL consistently improves the training efficiency of existing state-of-the-art (SOTA) CL methods by at most 23X less training FLOPs, and, surprisingly, further improves the SOTA accuracy by at most 1.7%. SparCL also outperforms competitive baselines obtained from adapting SOTA sparse training methods to the CL setting in both efficiency and accuracy. We also evaluate the effectiveness of SparCL on a real mobile phone, further indicating the practical potential of our method.