Optimal Service Elasticity in Large-Scale Distributed Systems

Optimal Service Elasticity in Large-Scale Distributed Systems
复制标题

大规模分布式系统中的最佳服务弹性

DOI:
--
复制
发表时间:
2017
期刊:
Proceedings of the ACM on Measurement and Analysis of Computing Systems
影响因子:
--
通讯作者:
J. V. Leeuwaarden
J. V. Leeuwaarden
中科院分区:
--
文献类型:
--
作者:
Debankur Mukherjee;Souvik Dhara;S. Borst;J. V. Leeuwaarden

文献摘要

被引文献

相似文献

大规模云网络和数据中心面临的一个根本挑战是实现高效的服务器利用率并限制能耗,同时在存在不确定和随时间变化的需求模式的情况下提供出色的用户感知性能。自动扩展提供了一种流行的范例,用于在满足性能目标的同时响应于需求自动调整服务容量,并且在文献中已经广泛研究了由业务驱动的自动扩展技术。然而,在典型的数据中心架构和云环境中,没有维护集中式队列,负载平衡算法立即在并行队列之间分配传入任务。在这些具有大量服务器的分布式设置中,集中式的业务驱动的自动缩放技术涉及大量的通信开销和主要的实现负担,或者甚至可能根本不可行。出于上述问题,我们提出了一个联合的自动缩放和负载均衡方案,不需要任何全局队列长度信息或系统参数的显式知识,但提供可证明的接近最优的服务弹性。我们建立了流体水平的动态拟议中的计划在一个政权的总交通量和名义服务能力增长大的比例。流体限制的结果表明,该方案实现了用户感知延迟性能以及能量消耗方面的渐近最优。具体来说,我们证明了任务的等待时间和空闲服务器消耗的相对能量部分都在极限内消失。同时,所提出的方案以分布式方式操作,并且每个任务仅涉及恒定的通信开销,从而确保大规模数据中心操作的可扩展性。大量的仿真实验证实了流体限制的结果,并表明,该方案可以匹配的用户性能和能源消耗的最先进的方法,充分利用集中式队列。
A fundamental challenge in large-scale cloud networks and data centers is to achieve highly efficient server utilization and limit energy consumption, while providing excellent user-perceived performance in the presence of uncertain and time-varying demand patterns. Auto-scaling provides a popular paradigm for automatically adjusting service capacity in response to demand while meeting performance targets, and queue-driven auto-scaling techniques have been widely investigated in the literature. In typical data center architectures and cloud environments however, no centralized queue is maintained, and load balancing algorithms immediately distribute incoming tasks among parallel queues. In these distributed settings with vast numbers of servers, centralized queue-driven auto-scaling techniques involve a substantial communication overhead and major implementation burden, or may not even be viable at all. Motivated by the above issues, we propose a joint auto-scaling and load balancing scheme which does not require any global queue length information or explicit knowledge of system parameters, and yet provides provably near-optimal service elasticity. We establish the fluid-level dynamics for the proposed scheme in a regime where the total traffic volume and nominal service capacity grow large in proportion. The fluid-limit results show that the proposed scheme achieves asymptotic optimality in terms of user-perceived delay performance as well as energy consumption. Specifically, we prove that both the waiting time of tasks and the relative energy portion consumed by idle servers vanish in the limit. At the same time, the proposed scheme operates in a distributed fashion and involves only constant communication overhead per task, thus ensuring scalability in massive data center operations. Extensive simulation experiments corroborate the fluid-limit results, and demonstrate that the proposed scheme can match the user performance and energy consumption of state-of-the-art approaches that do take full advantage of a centralized queue.