Sinan: ML-based and QoS-aware resource management for cloud microservices

Sinan: ML-based and QoS-aware resource management for cloud microservices
复制标题

DOI:
10.1145/3445814.3446693
复制
发表时间:
2021-04
期刊:
Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Yanqi Zhang;Weizhe Hua;Zhuangzhuang Zhou;Ed Suh;Christina Delimitrou;WeizheHua
Yanqi Zhang;Weizhe Hua;Zhuangzhuang Zhou;Ed Suh;Christina Delimitrou;WeizheHua
中科院分区:
其他
文献类型:
--
作者:
Yanqi Zhang;Weizhe Hua;Zhuangzhuang Zhou;Ed Suh;Christina Delimitrou;WeizheHua

文献摘要

被引文献

相似文献

云应用程序越来越多地从大型整体服务转移到大量松散耦合的专业微服务。尽管它们在促进发展,部署,模块化和隔离方面具有优势,但微服务使资源管理变得复杂,因为它们之间的依赖性引入了背压效应和级联违反QoS的侵犯。我们介绍了Sinan,这是一个由数据驱动的集群管理器,用于在线和QoS-Aware的交互式云微服务。 Sinan利用一组可扩展且经过验证的机器学习模型来确定微服务之间依赖关系的性能影响,并以维护端到端尾延迟目标的方式分配适当的资源。我们在用微服务(例如社交网络和酒店预订网站)上建造的代表性端到端应用程序上,在Google Compegute Engine(GCE)上对Sinan评估了Sinan。我们表明,与先前的工作相反,Sinan始终符合QoS,同时还保持群集利用率很高,这与先前的工作相反,这会导致不可预测的绩效或牺牲资源效率。此外,Sinan中的技术是可以解释的,这意味着云操作员可以从ML模型中产生有关如何更好地部署和设计其应用程序以降低无法预测的性能的见解。
Cloud applications are increasingly shifting from large monolithic services, to large numbers of loosely-coupled, specialized microservices. Despite their advantages in terms of facilitating development, deployment, modularity, and isolation, microservices complicate resource management, as dependencies between them introduce backpressure effects and cascading QoS violations. We present Sinan, a data-driven cluster manager for interactive cloud microservices that is online and QoS-aware. Sinan leverages a set of scalable and validated machine learning models to determine the performance impact of dependencies between microservices, and allocate appropriate resources per tier in a way that preserves the end-to-end tail latency target. We evaluate Sinan both on dedicated local clusters and large-scale deployments on Google Compute Engine (GCE) across representative end-to-end applications built with microservices, such as social networks and hotel reservation sites. We show that Sinan always meets QoS, while also maintaining cluster utilization high, in contrast to prior work which leads to unpredictable performance or sacrifices resource efficiency. Furthermore, the techniques in Sinan are explainable, meaning that cloud operators can yield insights from the ML models on how to better deploy and design their applications to reduce unpredictable performance.