IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency

IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
复制标题

DOI:
10.5070/sr34163500
复制
发表时间:
2023-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Saeid Ghafouri;Kamran Razavi;Mehran Salmani;Alireza Sanaee;T. Lorido-Botran;Lin Wang;Joseph Doyle;Pooyan Jamshidi
Saeid Ghafouri;Kamran Razavi;Mehran Salmani;Alireza Sanaee;T. Lorido-Botran;Lin Wang;Joseph Doyle;Pooyan Jamshidi
中科院分区:
其他
文献类型:
--
作者:
Saeid Ghafouri;Kamran Razavi;Mehran Salmani;Alireza Sanaee;T. Lorido-Botran;Lin Wang;Joseph Doyle;Pooyan Jamshidi

文献摘要

相似文献

鉴于其紧密的端到端延迟要求,有效地优化多模型推理管道以快速,准确和具有成本效益的推断是机器学习生产系统的关键挑战。为了简化探索延迟,准确性和推理管道成本的巨大而复杂的权衡空间,提供商经常选择考虑其中之一。但是,挑战在于协调延迟,准确性和成本权衡。为了应对这一挑战,并提出了一种解决推理管道中有效管理模型变体的解决方案,我们提出了IPA,这是一个在线深度学习推理管道适应系统,该系统有效地利用模型变体来为每个深度学习任务。模型变体是针对相同深度学习任务的预训练模型的不同版本,并且资源需求,延迟和准确性的变化。 IPA动态配置批次大小,复制和模型变体,以优化准确性,最小化成本并使用整数编程满足用户定义的延迟服务水平协议(SLA)。它支持多目标设置,以在准确性和成本目标之间实现不同的权衡,同时还可以适应不同的工作量和动态流量模式。与现有方法相比,导航更广泛的配置允许\ namex {}在成本和准确性目标之间实现更好的权衡。在具有五个现实世界推理管道的Kubernetes实施中进行的广泛实验表明,IPA可提高端到端的准确性多达21%,而成本提高最低。复制的代码和数据可在https://github.com/reconfigurable-ml-pipeline/ipa上获得。
Efficiently optimizing multi-model inference pipelines for fast, accurate, and cost-effective inference is a crucial challenge in machine learning production systems, given their tight end-to-end latency requirements. To simplify the exploration of the vast and intricate trade-off space of latency, accuracy, and cost in inference pipelines, providers frequently opt to consider one of them. However, the challenge lies in reconciling latency, accuracy, and cost trade-offs. To address this challenge and propose a solution to efficiently manage model variants in inference pipelines, we present IPA, an online deep learning Inference Pipeline Adaptation system that efficiently leverages model variants for each deep learning task. Model variants are different versions of pre-trained models for the same deep learning task with variations in resource requirements, latency, and accuracy. IPA dynamically configures batch size, replication, and model variants to optimize accuracy, minimize costs, and meet user-defined latency Service Level Agreements (SLAs) using Integer Programming. It supports multi-objective settings for achieving different trade-offs between accuracy and cost objectives while remaining adaptable to varying workloads and dynamic traffic patterns. Navigating a wider variety of configurations allows \namex{} to achieve better trade-offs between cost and accuracy objectives compared to existing methods. Extensive experiments in a Kubernetes implementation with five real-world inference pipelines demonstrate that IPA improves end-to-end accuracy by up to 21% with a minimal cost increase. The code and data for replications are available at https://github.com/reconfigurable-ml-pipeline/ipa.