Optimizing Inference Serving on Serverless Platforms

Optimizing Inference Serving on Serverless Platforms
复制标题

DOI:
10.14778/3547305.3547313
复制
发表时间:
2022-06
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Ahsan Ali;Riccardo Pinciroli;Feng Yan;E. Smirni
Ahsan Ali;Riccardo Pinciroli;Feng Yan;E. Smirni
中科院分区:
其他
文献类型:
--
作者:
Ahsan Ali;Riccardo Pinciroli;Feng Yan;E. Smirni

文献摘要

相似文献

无服务器计算因具有自动资源缩放、易于使用以及按使用付费的成本模式,在机器学习(ML)服务工作负载方面越来越受欢迎。现有的无服务器平台在基于图像的ML推理方面运行良好,因为这类请求在服务需求上是同质的。然而,自然语言处理的最新进展无法从现有无服务器平台中充分获益,因为其请求本质上是异质的。由于无服务器平台采用按使用付费的定价模式,对请求进行批处理可以显著提高ML服务效率,同时降低成本。但是,对异质的ML请求进行批处理会导致额外的计算开销,因为在同一批次内,小请求需要“填充”到与大请求相同的大小。做出有效的批处理决策(即哪些请求应该一起批处理以及为什么)并非易事:填充开销加上无服务器的自动缩放形成了一个复杂的优化问题。为了解决这个问题,我们开发了多缓冲区服务(MBS)框架,该框架对异质的ML推理服务请求的批处理进行优化,以在满足服务水平目标(SLOs)的同时将成本降至最低。MBS的核心是一个由贝叶斯优化器增强的分析模型驱动的性能和成本估算器。MBS在AWS上使用突发工作负载进行了原型设计和评估。实验结果表明,MBS在保持SLOs的同时,在成本节约方面比现有技术高出多达8倍,同时将填充开销降低多达37倍,无服务器函数调用次数减少3倍。
Serverless computing is gaining popularity for machine learning (ML) serving workload due to its autonomous resource scaling, easy to use and pay-per-use cost model. Existing serverless platforms work well for image-based ML inference, where requests are homogeneous in service demands. That said, recent advances in natural language processing could not fully benefit from existing serverless platforms as their requests are intrinsically heterogeneous. Batching requests for processing can significantly increase ML serving efficiency while reducing monetary cost, thanks to the pay-per-use pricing model adopted by serverless platforms. Yet, batching heterogeneous ML requests leads to additional computation overhead as small requests need to be "padded" to the same size as large requests within the same batch. Reaching effective batching decisions (i.e., which requests should be batched together and why) is non-trivial: the padding overhead coupled with the serverless auto-scaling forms a complex optimization problem. To address this, we develop Multi-Buffer Serving (MBS), a framework that optimizes the batching of heterogeneous ML inference serving requests to minimize their monetary cost while meeting their service level objectives (SLOs). The core of MBS is a performance and cost estimator driven by analytical models supercharged by a Bayesian optimizer. MBS is prototyped and evaluated on AWS using bursty workloads. Experimental results show that MBS preserves SLOs while outperforming the state-of-the-art by up to 8 x in terms of cost savings while minimizing the padding overhead by up to 37 x with 3 x less number of serverless function invocations.