FuncPipe: A Pipelined Serverless Framework for Fast and Cost-Efficient Training of Deep Learning Models

FuncPipe: A Pipelined Serverless Framework for Fast and Cost-Efficient Training of Deep Learning Models
复制标题

DOI:
10.1145/3570607
复制
发表时间:
2022-04
期刊:
Proceedings of the ACM on Measurement and Analysis of Computing Systems
影响因子:
--
通讯作者:
Yunzhuo Liu;Bo Jiang;Tian Guo;Zimeng Huang;Wen-ping Ma;Xinbing Wang;Chenghu Zhou
Yunzhuo Liu;Bo Jiang;Tian Guo;Zimeng Huang;Wen-ping Ma;Xinbing Wang;Chenghu Zhou
中科院分区:
其他
文献类型:
--
作者:
Yunzhuo Liu;Bo Jiang;Tian Guo;Zimeng Huang;Wen-ping Ma;Xinbing Wang;Chenghu Zhou

文献摘要

相似文献

在云端训练深度学习(DL)模型已成为一种常态。随着无服务器计算的出现以及其真正的按需付费定价和可扩展性的优势,系统研究人员最近已开始为基于无服务器的训练提供支持。然而,在无服务器平台上训练DL模型的能力受到当今无服务器基础设施的资源限制以及DL模型对内存和带宽的巨大需求的阻碍。本文介绍了FuncPipe,这是一种专为无服务器平台设计的新型流水线训练框架,能够实现DL模型的快速且低成本的训练。FuncPipe的设计关键在于可以利用模型划分来弥补无服务器函数的能力与DL训练需求之间在内存和带宽方面的差距。概念上虽然简单,但我们必须回答几个设计问题,包括如何划分模型、配置每个无服务器函数以及利用每个函数的上行/下行带宽。特别是,我们为无服务器环境定制了一种微批次调度策略,它是后续优化的基础。我们的混合整数二次规划公式自动且同时配置无服务器资源和划分模型以适应资源限制。最后,我们通过一种新型的流水线分散 - 归约算法提高了基于存储的同步的带宽效率。我们在两个流行的云无服务器平台上实现了FuncPipe,并表明与最先进的基于无服务器的框架相比,它实现了7% - 77%的成本节约以及1.3倍 - 2.2倍的加速。
Training deep learning (DL) models in the cloud has become a norm. With the emergence of serverless computing and its benefits of true pay-as-you-go pricing and scalability, systems researchers have recently started to provide support for serverless-based training. However, the ability to train DL models on serverless platforms is hindered by the resource limitations of today's serverless infrastructure and DL models' explosive requirement for memory and bandwidth. This paper describes FuncPipe, a novel pipelined training framework specifically designed for serverless platforms that enable fast and low-cost training of DL models. FuncPipe is designed with the key insight that model partitioning can be leveraged to bridge both memory and bandwidth gaps between the capacity of serverless functions and the requirement of DL training. Conceptually simple, we have to answer several design questions, including how to partition the model, configure each serverless function, and exploit each function's uplink/downlink bandwidth. In particular, we tailor a micro-batch scheduling policy for the serverless environment, which serves as the basis for the subsequent optimization. Our Mixed-Integer Quadratic Programming formulation automatically and simultaneously configures serverless resources and partitions models to fit within the resource constraints. Lastly, we improve the bandwidth efficiency of storage-based synchronization with a novel pipelined scatter-reduce algorithm. We implement FuncPipe on two popular cloud serverless platforms and show that it achieves 7%-77% cost savings and 1.3X-2.2X speedup compared to state-of-the-art serverless-based frameworks.