Model-driven Cluster Resource Management for AI Workloads in Edge Clouds

Model-driven Cluster Resource Management for AI Workloads in Edge Clouds
复制标题

DOI:
10.1145/3582080
复制
发表时间:
2022-01
影响因子:
2.7
通讯作者:
Qianlin Liang;Walid A. Hanafy;Ahmed Ali-Eldin;Prashant Shenoy
Qianlin Liang;Walid A. Hanafy;Ahmed Ali-Eldin;Prashant Shenoy
中科院分区:
计算机科学4区
文献类型:
--
作者:
Qianlin Liang;Walid A. Hanafy;Ahmed Ali-Eldin;Prashant Shenoy

文献摘要

被引文献

相似文献

由于物联网(IoT)分析和增强现实等新兴边缘应用具有严格的延迟限制,最近提出了硬件AI加速器来加速这些应用程序运行的深度神经网络(DNN)推理。资源受限的边缘服务器和加速器往往会在多个物联网应用程序中进行复用,这可能会在延迟敏感的工作负载之间引入性能干扰。在本文中,我们设计了分析模型来捕获DNN推理工作负载在不同复用和并发行为下在共享边缘加速器(如GPU和edgeTPU)上的性能。在使用广泛的实验验证了我们的模型之后,我们使用它们来设计各种集群资源管理算法,以智能地管理边缘加速器上的多个应用程序,同时尊重它们的延迟限制。我们在Kubernetes中实现了我们的系统原型,并表明与传统的背包托管算法相比,我们的系统可以在异构多租户边缘集群中托管2.3倍以上的DNN应用程序,并且没有延迟违规。
Since emerging edge applications such as Internet of Things (IoT) analytics and augmented reality have tight latency constraints, hardware AI accelerators have been recently proposed to speed up deep neural network (DNN) inference run by these applications. Resource-constrained edge servers and accelerators tend to be multiplexed across multiple IoT applications, introducing the potential for performance interference between latency-sensitive workloads. In this article, we design analytic models to capture the performance of DNN inference workloads on shared edge accelerators, such as GPU and edgeTPU, under different multiplexing and concurrency behaviors. After validating our models using extensive experiments, we use them to design various cluster resource management algorithms to intelligently manage multiple applications on edge accelerators while respecting their latency constraints. We implement a prototype of our system in Kubernetes and show that our system can host 2.3× more DNN applications in heterogeneous multi-tenant edge clusters with no latency violations when compared to traditional knapsack hosting algorithms.