WattWiser: Power & Resource-Efficient Scheduling for Multi-Model Multi-GPU Inference Servers

WattWiser: Power & Resource-Efficient Scheduling for Multi-Model Multi-GPU Inference Servers
复制标题

WattWiser:电源

DOI:
10.1145/3634769.3634807
复制
发表时间:
2023
期刊:
Proceedings of the 14th International Green and Sustainable Computing Conference
影响因子:
--
通讯作者:
Daniel Wong
Daniel Wong
中科院分区:
--
文献类型:
--
作者:
Ali Jahanshahi;Mohammadreza Rezvani;Daniel Wong

文献摘要

参考文献

相似文献

随着机器学习(ML)应用越来越多地集成到云服务中,提供高吞吐量的机器学习推理服务已成为云服务提供商的主要需求。推理请求需要以有限的延迟响应每个请求,以保持一致的服务级别目标(SLO)。为了确保SLO,推理服务器配备了多个GPU以满足计算需求。然而,多GPU系统非常耗电。为了解决这个问题,理想的做法是将负载整合到GPU的子集,并可能共享GPU,以最大限度地降低功耗,而不违反SLO。通过整合GPU和潜在的共享GPU,我们可以降低多GPU推理服务器的功耗。然而,多个推理模型通常共享相同的推理服务器,这在多模型多GPU推理服务器环境中增加了显著的挑战。在本文中,我们将探讨这在实现电源效率方面所带来的挑战。我们介绍WattWiser,这是一种模型管理和调度策略,可在共享GPU的多模型环境中实现节能。我们的研究结果表明,WattWiser可以降低34%的功耗,同时服务于多个型号并保持SLO。
With the increasing integration of Machine Learning (ML) applications into cloud services, providing high throughput Machine Learning inference serving has become a major demand for cloud service providers. The inference requests need to respond with bounded latency for each request to maintain a consistent Service-Level Objective (SLO). To ensure SLO, inference servers are equipped with multiple GPUs to satisfy the computational requirements. However, multi-GPU systems are extremely power-hungry. To resolve this, it is ideal to consolidate the load to a sub-set of GPUs, and potentially share GPUs, in order to minimize power consumption, without violating SLO. By consolidating GPUs and potentially sharing GPUs we can reduce the power consumption of multi-GPU inference servers. However, multiple inference models typically share the same inference server, which adds significant challenges in multi-model multi-GPU inference server environments. In this paper, we explore the challenges that this brings in achieving power efficiency. We introduce WattWiser, a model management and scheduling policy that achieves power savings in multi-model environments where GPUs are shared. Our results show that WattWiser can reduce power consumption by 34% while serving multiple models and maintaining the SLO.
DOI: 10.1145/3419111.3421284
发表时间: 2020-10
期刊: Proceedings of the 11th ACM Symposium on Cloud Computing
影响因子: --
作者:
Aditya Dhakal;Sameer G. Kulkarni;K. Ramakrishnan
通讯作者: Aditya Dhakal;Sameer G. Kulkarni;K. Ramakrishnan
DOI: 10.1109/hpca56546.2023.10071121
发表时间: 2023-02
期刊: 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子: --
作者:
M. Chow;Ali Jahanshahi;Daniel Wong
通讯作者: M. Chow;Ali Jahanshahi;Daniel Wong