Reconciling High Accuracy, Cost-Efficiency, and Low Latency of Inference Serving Systems

Reconciling High Accuracy, Cost-Efficiency, and Low Latency of Inference Serving Systems
复制标题

DOI:
10.1145/3578356.3592578
复制
发表时间:
2023-04
期刊:
Proceedings of the 3rd Workshop on Machine Learning and Systems
影响因子:
--
通讯作者:
Mehran Salmani;Saeid Ghafouri;Alireza Sanaee;Kamran Razavi;M. Muhlhauser;Joseph Doyle;Pooyan Jamshidi;Mohsen Sharif Iran University of Science;Technology;Queen Mary University London;Technical University of Darmstadt;U. O. N. Carolina
Mehran Salmani;Saeid Ghafouri;Alireza Sanaee;Kamran Razavi;M. Muhlhauser;Joseph Doyle;Pooyan Jamshidi;Mohsen Sharif Iran University of Science;Technology;Queen Mary University London;Technical University of Darmstadt;U. O. N. Carolina
中科院分区:
其他
文献类型:
--
作者:
Mehran Salmani;Saeid Ghafouri;Alireza Sanaee;Kamran Razavi;M. Muhlhauser;Joseph Doyle;Pooyan Jamshidi;Mohsen Sharif Iran University of Science;Technology;Queen Mary University London;Technical University of Darmstadt;U. O. N. Carolina

文献摘要

相似文献

在各种应用程序上使用机器学习(ML)推断正在急剧增长。 ML推理服务直接与用户互动,需要快速准确的响应。此外,这些服务面临着动态的请求工作负载,并在其计算资源中施加更改。无法右尺寸的计算资源导致延迟服务级别目标(SLO)违规或浪费计算资源。考虑到所有准确性,延迟和资源成本的支柱,适应动态工作负载都是具有挑战性的。为了应对这些挑战,我们提出了Indradapter,该Indradapter主动选择了一组ML模型变体及其资源分配以满足延迟SLO,同时最大程度地提高了由精度和成本组成的目标函数。与受欢迎的行业自动制剂(Kubernetes垂直POD Autoscaler)相比,Indradapter降低了SLO违规,成本高达65%和33%。
The use of machine learning (ML) inference for various applications is growing drastically. ML inference services engage with users directly, requiring fast and accurate responses. Moreover, these services face dynamic workloads of requests, imposing changes in their computing resources. Failing to right-size computing resources results in either latency service level objectives (SLOs) violations or wasted computing resources. Adapting to dynamic workloads considering all the pillars of accuracy, latency, and resource cost is challenging. In response to these challenges, we propose InfAdapter, which proactively selects a set of ML model variants with their resource allocations to meet latency SLO while maximizing an objective function composed of accuracy and cost. InfAdapter decreases SLO violation and costs up to 65% and 33%, respectively, compared to a popular industry autoscaler (Kubernetes Vertical Pod Autoscaler).