Holistic energy and failure aware workload scheduling in Cloud datacenters

Holistic energy and failure aware workload scheduling in Cloud datacenters
复制标题

DOI:
10.1016/j.future.2017.07.044
复制
发表时间:
2018
期刊:
Future Gener. Comput. Syst.
影响因子:
--
通讯作者:
Xiang Li;Xiaohong Jiang;Peter Garraghan;Zhaohui Wu
Xiang Li;Xiaohong Jiang;Peter Garraghan;Zhaohui Wu
中科院分区:
其他
文献类型:
--
作者:
Xiang Li;Xiaohong Jiang;Peter Garraghan;Zhaohui Wu

文献摘要

被引文献

相似文献

云计算在全球范围内的普及引起了学术界和工业界越来越多的兴趣,导致大规模和复杂的分布式系统的形成。这导致计算系统内的故障发生率增加,从而对用户感知的系统性能和任务可靠性产生重大负面影响。此类系统还消耗大量电力,导致提供商认为运营成本很高。虚拟化是云数据中心内普遍部署的技术,可以实现虚拟机的灵活调度,以最大限度地提高系统可靠性和能源效率。然而,现有的工作分别解决了这两个目标,对研究可靠且节能的计算基础设施的明确权衡提供了有限的理解。在本文中,我们提出了两种故障感知的节能调度算法,该算法利用了云数据中心的整体运行特征,包括冷却单元、计算基础设施和服务器故障。通过对云数据中心的电源和故障概况进行全面建模,我们提出了工作负载调度算法 Ella-WandElla-B,能够减少冷却和计算能耗,同时最大限度地减少系统故障的影响。提出了一种新颖的整体指标,结合能源效率和可靠性来指定各种算法的性能。我们在各种故障预测精度和工作负载强度的系统条件下针对Random、MaxUtil、TASA、MTTE 和OBFIT 评估我们的算法。评估结果表明,Ella-W可以减少能源消耗29.5%,任务完成率提高3.6%,而Ella-Bruce能源消耗减少32.7%,任务完成率没有下降。
The global uptake of Cloud computing has attracted increased interest within both academia and industry resulting in the formation of large-scale and complex distributed systems. This has led to increased failure occurrence within computing systems that induce substantial negative impact upon system performance and task reliability perceived by users. Such systems also consume vast quantities of power, resulting in significant operational costs perceived by providers. Virtualization – a commonly deployed technology within Cloud datacenters – can enable flexible scheduling of virtual machines to maximize system reliability and energy-efficiency. However, existing work address these two objectives separately, providing limited understanding towards studying the explicit trade-offs towards dependable and energy-efficient compute infrastructure. In this paper, we propose two failure-aware energy-efficient scheduling algorithms that exploit the holistic operational characteristics of the Cloud datacenter comprising the cooling unit, computing infrastructure and server failures. By comprehensively modeling the power and failure profiles of a Cloud datacenter, we propose workload scheduling algorithmsElla-WandElla-B, capable of reducing cooling and compute energy while minimizing the impact of system failures. A novel and overall metric is proposed that combines energy efficiency and reliability to specify the performance of various algorithms. We evaluate our algorithms againstRandom,MaxUtil,TASA,MTTEandOBFITunder various system conditions of failure prediction accuracy and workload intensity. Evaluation results demonstrate thatElla-Wcan reduce energy usage by 29.5% and improve task completion rate by 3.6%, whileElla-Breduces energy usage by 32.7% with no degradation to task completion rate.