Performance Implications of Failures in Large-Scale Cluster Scheduling

Performance Implications of Failures in Large-Scale Cluster Scheduling
复制标题

DOI:
10.1007/11407522_13
复制
发表时间:
2004-06
期刊:
--
影响因子:
--
通讯作者:
Yanyong Zhang;M. Squillante;A. Sivasubramaniam;R. Sahoo
Yanyong Zhang;M. Squillante;A. Sivasubramaniam;R. Sahoo
中科院分区:
其他
文献类型:
--
作者:
Yanyong Zhang;M. Squillante;A. Sivasubramaniam;R. Sahoo

文献摘要

被引文献

相似文献

随着我们不断发展成为大规模并行系统,其中许多采用数百个计算引擎承担关键任务的角色,这是至关重要的设计这些系统预测和适应故障的发生。故障成为这种大规模系统的常见特征,人们不能继续将其视为例外。尽管目前和越来越多的重要性,在这些系统中的故障,我们的理解这些关键问题对并行计算环境的性能影响是非常有限的。在本文中,我们开发了一个通用的故障建模框架的基础上,最近的结果从大规模集群,然后我们利用这个框架进行详细的性能分析故障对系统性能的影响范围广泛的调度策略。我们的研究结果表明,这样的故障可以有显着的影响,平均作业响应时间和平均作业减速现有的调度策略,忽略故障。因此,我们研究不同的调度机制和政策,以解决这些性能问题。我们的研究结果表明,定期检查点的工作似乎没有做什么,以缓解这个问题。另一方面,我们证明了故障发生的空间和时间相关性的信息可以是非常有用的设计调度(作业分配)策略,以提高系统性能,前者提供了最大的好处。
As we continue to evolve into large-scale parallel systems, many of them employing hundreds of computing engines to take on mission-critical roles, it is crucial to design those systems anticipating and accommodating the occurrence of failures. Failures become a commonplace feature of such large-scale systems, and one cannot continue to treat them as an exception. Despite the current and increasing importance of failures in these systems, our understanding of the performance impact of these critical issues on parallel computing environments is extremely limited. In this paper we develop a general failure modeling framework based on recent results from large-scale clusters and then we exploit this framework to conduct a detailed performance analysis of the impact of failures on system performance for a wide range of scheduling policies. Our results demonstrate that such failures can have a significant impact on the mean job response time and mean job slowdown under existing scheduling policies that ignore failures. We therefore investigate different scheduling mechanisms and policies to address these performance issues. Our results show that periodic checkpointing of jobs seems to do little to ease this problem. On the other hand, we demonstrate that information about the spatial and temporal correlation of failure occurrences can be very useful in designing a scheduling (job allocation) strategy to enhance system performance, with the former providing the greatest benefits.