QuMan

QuMan
复制标题

DOI:
10.1145/3210560
复制
发表时间:
2018-08
期刊:
ACM Transactions on Architecture and Code Optimization (TACO)
影响因子:
--
通讯作者:
Yannis Sfakianakis;C. Kozanitis;Christos Kozyrakis;A. Bilas
Yannis Sfakianakis;C. Kozanitis;Christos Kozyrakis;A. Bilas
中科院分区:
其他
文献类型:
--
作者:
Yannis Sfakianakis;C. Kozanitis;Christos Kozyrakis;A. Bilas

文献摘要

被引文献

相似文献

现代数据中心整合工作负载以提高服务器利用率、降低总拥有成本并科普扩展限制。然而,服务器资源共享会在应用程序之间引入性能干扰,从而增加性能波动,这会对用户体验产生负面影响。因此,一个具有挑战性的问题是提高服务器的利用率,同时保持应用程序的QoS。在本文中,我们将介绍一个服务器资源管理器QuMan,它使用应用程序隔离和分析来提高服务器利用率,同时控制应用程序QoS的降低。以前的解决方案要么估计跨应用的干扰,然后将托管限制为“兼容”应用,要么假设应用要求是已知的。相反,QuMan估计应用程序所需的资源。它使用隔离机制为应用程序创建适当大小的资源片,并任意地合并应用程序。QuMan的机制可以用于各种准入控制政策,我们探讨了两个这样的政策的潜力:(1)一个政策,允许用户指定一个最低的性能阈值和(2)一个自动化的政策,它的操作没有用户输入,是基于一个新的组合QoS利用率度量。我们在Linux服务器上实现了QuMan,并使用容器和真实的应用程序评估了其有效性。我们的单节点结果表明,QuMan非常有效地平衡了服务器利用率和应用程序性能之间的权衡,因为它实现了80%的服务器利用率,而每个应用程序的性能不会低于各自独立性能的80%。我们还将QuMan部署在一个由100个AWS实例组成的集群上,这些实例由Sparrow调度程序的修改版本管理[37],并且我们观察到,与由本地Sparrow或Apache Mesos管理时相同负载下的相同集群的性能相比,在高度利用的集群上的应用程序性能提高了48%。
Modern data centers consolidate workloads to increase server utilization and reduce total cost of ownership, and cope with scaling limitations. However, server resource sharing introduces performance interference across applications and, consequently, increases performance volatility, which negatively affects user experience. Thus, a challenging problem is to increase server utilization while maintaining application QoS. In this article, we present QuMan, a server resource manager that uses application isolation and profiling to increase server utilization while controlling degradation of application QoS. Previous solutions, either estimate interference across applications and then restrict colocation to “compatible” applications, or assume that application requirements are known. Instead, QuMan estimates the required resources of applications. It uses an isolation mechanism to create properly-sized resource slices for applications, and arbitrarily colocates applications. QuMan’s mechanisms can be used with a variety of admission control policies, and we explore the potential of two such policies: (1) A policy that allows users to specify a minimum performance threshold and (2) an automated policy, which operates without user input and is based on a new combined QoS-utilization metric. We implement QuMan on top of Linux servers, and we evaluate its effectiveness using containers and real applications. Our single-node results show that QuMan balances highly effectively the tradeoff between server utilization and application performance, as it achieves 80% server utilization while the performance of each application does not drop below 80% the respective standalone performance. We also deploy QuMan on a cluster of 100 AWS instances that are managed by a modified version of the Sparrow scheduler [37] and, we observe a 48% increase in application performance on a highly utilized cluster, compared to the performance of the same cluster under the same load when it is managed by native Sparrow or Apache Mesos.