State Monitoring in Cloud Datacenters

State Monitoring in Cloud Datacenters
复制标题

DOI:
10.1109/tkde.2011.70
复制
发表时间:
2011-09
影响因子:
8.9
通讯作者:
S. Meng;Ling Liu;Ting Wang
S. Meng;Ling Liu;Ting Wang
中科院分区:
计算机科学2区
文献类型:
--
作者:
S. Meng;Ling Liu;Ting Wang

文献摘要

被引文献

相似文献

监控分布式云应用程序的全局状态是云数据中心管理的一项关键功能。状态监控需要满足两个苛刻的目标:高水平的正确性,确保零错误率或低错误率;以及高通信效率,这在检测状态更新时需要最小的通信成本。大多数现有工作都遵循即时模型,只要违反约束就会触发状态警报。该模型可能会因瞬时值爆发和异常值而导致频繁且不必要的警报。针对此类警报的对策可能会进一步导致操作出现问题。在本文中,我们提出了一种基于 WIndow 的 StatE 监控 (WISE) 框架,用于有效管理云应用程序。基于窗口的状态监控报告仅当状态违规在时间窗口内连续发生时才会发出警报。我们证明,它不仅对价值爆发和异常值更具弹性,而且基于四项技术贡献以分布式方式实施时能够节省大量通信。首先,我们提出了基于窗口的状态监控和集中参数调整的架构设计和部署选项。其次,我们开发了一种新的分布式参数调整方案,使 WISE 能够扩展到更多的监控节点,因为每个节点都可以在没有全局信息的情况下反应性地调整其监控参数。第三,我们介绍了两种优化技术,包括它们的设计原理、正确性和使用模型,以进一步降低通信成本。最后,我们对WISE的可扩展性进行了深入的实证研究,并评估了分布式调优方案和两次性能优化带来的改进。我们的结果表明,与即时监控方法相比,WISE 减少了 50-90% 的通信,并且改进后的 WISE 比其集中式版本获得了明显的可扩展性优势。
Monitoring global states of a distributed cloud application is a critical functionality for cloud datacenter management. State monitoring requires meeting two demanding objectives: high level of correctness, which ensures zero or low error rate, and high communication efficiency, which demands minimal communication cost in detecting state updates. Most existing work follows an instantaneous model which triggers state alerts whenever a constraint is violated. This model may cause frequent and unnecessary alerts due to momentary value bursts and outliers. Countermeasures of such alerts may further cause problematic operations. In this paper, we present a WIndow-based StatE monitoring (WISE) framework for efficiently managing cloud applications. Window-based state monitoring reports alerts only when state violation is continuous within a time window. We show that it is not only more resilient to value bursts and outliers, but also able to save considerable communication when implemented in a distributed manner based on four technical contributions. First, we present the architectural design and deployment options for window-based state monitoring with centralized parameter tuning. Second, we develop a new distributed parameter tuning scheme enabling WISE to scale to much more monitoring nodes as each node tunes its monitoring parameters reactively without global information. Third, we introduce two optimization techniques, including their design rationale, correctness and usage model, to further reduce the communication cost. Finally, we provide an in-depth empirical study of the scalability of WISE, and evaluate the improvement brought by the distributed tuning scheme and the two performance optimizations. Our results show that WISE reduces communication by 50-90 percent compared with instantaneous monitoring approaches, and the improved WISE gains a clear scalability advantage over its centralized version.