BATCH: Machine Learning Inference Serving on Serverless Platforms with Adaptive Batching

BATCH: Machine Learning Inference Serving on Serverless Platforms with Adaptive Batching
复制标题

DOI:
10.1109/sc41405.2020.00073
复制
发表时间:
2020-11
期刊:
SC20: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Ahsan Ali;Riccardo Pinciroli;Feng Yan;E. Smirni
Ahsan Ali;Riccardo Pinciroli;Feng Yan;E. Smirni
中科院分区:
其他
文献类型:
--
作者:
Ahsan Ali;Riccardo Pinciroli;Feng Yan;E. Smirni

文献摘要

相似文献

无服务器计算是一种新的按使用付费的云服务范式,它能自动为无状态函数进行资源扩展,并有可能促进突发的机器学习服务。批处理对于机器学习推理的延迟性能和成本效益至关重要,但不幸的是,由于现有无服务器平台的无状态设计,它们不支持批处理。我们的实验表明,如果没有批处理,机器学习服务无法获得无服务器计算的优势。在本文中,我们提出了BATCH,一个用于在无服务器平台上支持高效机器学习服务的框架。BATCH使用一个优化器来提供推理尾延迟保证和成本优化,并实现自适应批处理支持。我们在AWS Lambda和流行的机器学习推理系统上构建了BATCH的原型。评估验证了分析优化器的准确性,并展示了相对于最先进的方法MArk和实际应用中的工具SageMaker在性能和成本方面的优势。
Serverless computing is a new pay-per-use cloud service paradigm that automates resource scaling for stateless functions and can potentially facilitate bursty machine learning serving. Batching is critical for latency performance and cost-effectiveness of machine learning inference, but unfortunately it is not supported by existing serverless platforms due to their stateless design. Our experiments show that without batching, machine learning serving cannot reap the benefits of serverless computing. In this paper, we present BATCH, a framework for supporting efficient machine learning serving on serverless platforms. BATCH uses an optimizer to provide inference tail latency guarantees and cost optimization and to enable adaptive batching support. We prototype BATCH atop of AWS Lambda and popular machine learning inference systems. The evaluation verifies the accuracy of the analytic optimizer and demonstrates performance and cost advantages over the state-of-the-art method MArk and the state-of-the-practice tool SageMaker.