Egeria: Efficient DNN Training with Knowledge-Guided Layer Freezing

Egeria: Efficient DNN Training with Knowledge-Guided Layer Freezing
复制标题

DOI:
10.1145/3552326.3587451
复制
发表时间:
2022-01
期刊:
Proceedings of the Eighteenth European Conference on Computer Systems
影响因子:
--
通讯作者:
Yiding Wang;D. Sun;Kai Chen-;Fan Lai;Mosharaf Chowdhury
Yiding Wang;D. Sun;Kai Chen-;Fan Lai;Mosharaf Chowdhury
中科院分区:
其他
文献类型:
--
作者:
Yiding Wang;D. Sun;Kai Chen-;Fan Lai;Mosharaf Chowdhury

文献摘要

相似文献

训练深度神经网络(DNN)非常耗时。虽然大多数现有解决方案试图重叠/调度计算和通信以进行高效训练,但本文更进一步,通过DNN层冻结跳过计算和通信。我们的关键见解是,内部DNN层的训练进度差异很大,前端层通常比深层更早地训练好。为了探索这一点,我们首先引入训练可塑性的概念来量化内部DNN层的训练进度。然后,我们设计了Egeria,这是一个知识引导的DNN训练系统,它利用参考模型中的语义知识来准确评估各个层的训练可塑性,并安全地冻结收敛的层,从而节省了相应的向后计算和通信。我们的参考模型是使用量化技术动态生成的,并在可用CPU上异步运行向前操作,以最大限度地减少开销。此外,Egeria通过预取缓存冻结层的中间输出,以进一步跳过向前计算。我们的实现和测试平台实验与流行的视觉和语言模型表明,Egeria实现了19%-43%的训练加速比w.r.t.在不牺牲准确性的前提下
Training deep neural networks (DNNs) is time-consuming. While most existing solutions try to overlap/schedule computation and communication for efficient training, this paper goes one step further by skipping computing and communication through DNN layer freezing. Our key insight is that the training progress of internal DNN layers differs significantly, and front layers often become well-trained much earlier than deep layers. To explore this, we first introduce the notion of training plasticity to quantify the training progress of internal DNN layers. Then we design Egeria, a knowledge-guided DNN training system that employs semantic knowledge from a reference model to accurately evaluate individual layers' training plasticity and safely freeze the converged ones, saving their corresponding backward computation and communication. Our reference model is generated on the fly using quantization techniques and runs forward operations asynchronously on available CPUs to minimize the overhead. In addition, Egeria caches the intermediate outputs of the frozen layers with prefetching to further skip the forward computation. Our implementation and testbed experiments with popular vision and language models show that Egeria achieves 19%-43% training speedup w.r.t. the state-of-the-art without sacrificing accuracy.