ModelKeeper: Accelerating DNN Training via Automated Training Warmup

ModelKeeper: Accelerating DNN Training via Automated Training Warmup
复制标题

DOI:
--
复制
发表时间:
2023
影响因子:
15
通讯作者:
Fan Lai;Yinwei Dai;H. Madhyastha;Mosharaf Chowdhury
Fan Lai;Yinwei Dai;H. Madhyastha;Mosharaf Chowdhury
中科院分区:
化学1区
文献类型:
--
作者:
Fan Lai;Yinwei Dai;H. Madhyastha;Mosharaf Chowdhury

文献摘要

相似文献

随着机器学习(ML)模型的部署越来越多,ML开发人员正在训练或重新训练越来越多的深度神经网络(dnn)。他们这样做是为了找到最合适的模型,在满足目标环境的资源和时效性约束的同时满足他们的准确性要求。在大型共享集群中,越来越多的神经架构搜索(NAS)和训练工作通常会导致模型与来自相同或不同ML开发人员的其他模型共享架构相似性。然而,现有的解决方案并没有提供识别和利用这种相似性的系统机制。我们提出了ModelKeeper,这是第一个自动训练预热系统,通过在共享集群中重新利用以前训练过的模型来加速DNN训练。我们的关键见解是,通过转换已经训练好的模型的权重来初始化训练任务的模型,可以启动它并减少所需的训练总量。然而,随着时间的推移,提交的模型在其架构和准确性方面可能有所不同。给定要训练的新模型,ModelKeeper可扩展地识别其与先前训练的模型的体系结构相似性,选择相似度高且模型精度好的父模型,并在新模型权值预热期间执行结构感知的权值转换,以最大限度地保留来自父模型的信息。我们对数千个CV和NLP模型的评估表明,ModelKeeper的训练完成速度提高了1.3 × -4.3 ×,开销很小,模型精度没有降低。
With growing deployment of machine learning (ML) models, ML developers are training or re-training increasingly more deep neural networks (DNNs). They do so to find the most suitable model that meets their accuracy requirement while satisfying the resource and timeliness constraints of the target environment. In large shared clusters, the growing number of neural architecture search (NAS) and training jobs often result in models sharing architectural similarities with others from the same or a different ML developer. However, existing solutions do not provide a systematic mechanism to identify and leverage such similarities. We present ModelKeeper, the first automated training warmup system that accelerates DNN training by repurposing previously-trained models in a shared cluster. Our key insight is that initializing a training job’s model by transforming an already-trained model’s weights can jump-start it and reduce the total amount of training needed. However, models submitted over time can differ in their architectures and accuracy. Given a new model to train, ModelKeeper scalably identifies its architectural similarity with previously trained models, selects a parent model with high similarity and good model accuracy, and performs structure-aware transformation of weights to preserve maximal information from the parent model during the warmup of new model weights. Our evaluations across thousands of CV and NLP models show that ModelKeeper achieves 1.3 × –4.3 × faster training completion with little overhead and no reduction in model accuracy.