Improving Gang Scheduling through job performance analysis and malleability

Improving Gang Scheduling through job performance analysis and malleability
复制标题

通过工作绩效分析和可塑性改进班组调度

DOI:
10.1145/377792.377852
复制
发表时间:
2001
期刊:
影响因子:
2.2
通讯作者:
J. Labarta
J. Labarta
中科院分区:
计算机科学4区
文献类型:
--
作者:
J. Corbalán;X. Martorell;J. Labarta

文献摘要

被引文献

相似文献

OpenMP 编程模型为并行应用程序提供了一个非常重要的功能:作业延展性。作业延展性是应用程序动态调整其并行性以适应分配给它的处理器数量的能力。我们相信,工作的可延展性为应用程序提供了系统实现其最大性能所需的灵活性。我们还认为,系统不仅必须根据用户需求做出决策,还必须根据运行时性能测量来做出决策,以确保资源的有效利用。作业延展性是使运行时性能分析成为可能的应用程序特征。如果没有可延展性,应用程序将无法使其并行性适应系统决策。为了支持这些想法,我们提出了两种新方法来解决组调度的两个主要问题:时隙数量过多和碎片。我们的第一个建议是在 Gang Scheduling 的每个时隙内应用调度策略,以考虑到基于运行时测量计算的效率,在应用程序之间分配处理器。我们将此策略称为“性能驱动的组调度”。我们的第二种方法是一种新的重新打包算法,即 Compress&Join,它利用了作业的可延展性。该算法修改正在运行的应用程序的处理器分配,以使其适应系统需求,并最大限度地减少碎片和时隙数量。这些建议已在具有 64 个处理器的 SGI Origin 2000 中实施。结果显示了两者的有效性和便利性,考虑在运行时计算的作业性能分析来决定处理器分配,并使用灵活的编程模型使应用程序适应系统决策。
The OpenMP programming model provides parallel applications a very important feature: job malleability. Job malleability is the capacity of an application to dynamically adapt its parallelism to the number of processors allocated to it. We believe that job malleability provides to applications the flexibility that a system needs to achieve its maximum performance. We also defend that a system has to take its decisions not only based on user requirements but also based on run-time performance measurements to ensure the efficient use of resources. Job malleability is the application characteristic that makes possible the run-time performance analysis. Without malleability applications would not be able to adapt their parallelism to the system decisions. To support these ideas, we present two new approaches to attack the two main problems of Gang Scheduling: the excessive number of time slots and the fragmentation. Our first proposal is to apply a scheduling policy inside each time slot of Gang Scheduling to distribute processors among applications considering their efficiency, calculated based on run-time measurements. We call this policy Performance-Driven Gang Scheduling. Our second approach is a new re-packing algorithm, Compress&Join, that exploits the job malleability. This algorithm modifies the processor allocation of running applications to adapt it to the system necessities and minimize the fragmentation and number of time slots. These proposals have been implemented in a SGI Origin 2000 with 64 processors. Results show the validity and convenience of both, to consider the job performance analysis calculated at run-time to decide the processor allocation, and to use a flexible programming model that adapts applications to system decisions.