课题基金 / 基金详情

A Framework for Next-Generation Scheduling and Task Management for Extreme-Scale Computing

A Framework for Next-Generation Scheduling and Task Management for Extreme-Scale Computing
超大规模计算的下一代调度和任务管理框架
批准号:
0444417
负责人:
Kang Shin
金额:
$0.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-11-01 至 2009-10-31

项目摘要

项目成果

Kang Shin的其他基金

相似基金

相关文献

中文摘要
翻译
用于极端规模计算的下一代调度和任务管理框架kang G. Shin和Abhijit bose密歇根大学该研究项目的目标是开发用于支持极端规模计算的调度和资源管理的高效算法和强大软件。这样的计算可能需要成千上万的处理器配置成一个高端的计算系统。需要比当前使用的更有效的调度和资源管理系统来解决此类大型系统的可伸缩性和容错问题。同样的论点也适用于许多新兴应用程序日益复杂的需求。当前一代调度器主要管理为应用程序静态配置的cpu集群。应用程序被分配了固定数量的处理器或节点(对于SMP系统),并且在执行期间不需要修改其资源需求。这可能导致较低的资源利用率。此外,大多数当前调度器不为运行的工作负载提供透明的容错。一个或多个处理器/节点故障通常会终止这些处理器上当前正在执行的任务,从而导致浪费周期。容错调度对于未来HEC系统的可扩展性至关重要,因为它可以容纳大量的处理器。计算密集型和数据密集型应用越来越多地使用高端系统,这要求未来的调度系统在将任务映射到适当的资源时,必须解决协同调度cpu、I/O和网络资源的问题。在一些HEC体系结构中,由于底层互连拓扑,进程的放置对应用程序的整体性能有影响。协同调度和工作负载感知调度对于未来的高端计算系统都很重要。将开发一个集成的软件框架,用于极端规模计算系统的调度和资源管理,提供以下能力:(i)在线工作负载特征,(ii)基于时间序列建模和预测资源利用率和排队工作负载的预测调度,(iii)被一个或多个故障中断的应用程序的透明容错调度,以及(iv)将工作负载和资源特征和预测作为调度决策一部分的高效启发式和进化算法。这项工作代表了调度理论和健壮的软件开发的混合。这些算法和软件框架在密歇根大学高级计算中心(CAC)的HEC生产环境中。CAC HEC设施提供了一个由1400多个cpu组成的测试平台,代表多个处理器系列(AMD athlon, Opterons, Apple Xserve/G5)和几个互连系统,如千兆以太网和Myrinet。本研究的智力价值将是推动HEC系统调度和资源管理的最新技术。通过在CAC上实现和部署建议的框架,我们将能够从各种最终用户和应用程序中收集实际的工作负载跟踪。本研究将为此类系统的鲁棒容错调度算法及其实现软件的开发提供催化剂。我们特别解决了容错机制的可伸缩性,比如检查点/重新启动和I/O,它们可以扩展到数千个处理器。此外,我们提出的研究、推广和教育活动的整合将对其他HEC中心和调度研究界产生更广泛的影响。作为该项目的一部分开发的框架和我们的研究成果将通过开源软件和高质量出版物传播给工业界和学术界。
英文摘要
A Framework for Next-generation Scheduling and Task Management for Extreme-scale ComputingKang G. Shin and Abhijit BoseThe University of MichiganThe objective of this research project is to develop efficient algorithms and robust software for scheduling and resource management in support of extreme-scale computing. Such computations may involve tens of thousands of processors configured as a high-end computing system. More efficient scheduling and resource management systems than those currently used are needed to address scalability and fault-tolerance for such large systems. The same argument also holds for the increasingly complex requirements of many emerging applications. The current-generation schedulers primarily manage a cluster of CPUs that are configured statically for an application. An application is allocated a fixed number of processors or nodes (for SMP systems), and is not expected to modify its resource requirement during the execution. This may result in lower resource utilization. Furthermore, most of the current schedulers do not provide transparent fault-tolerance to the running workload. One or more processor/node failures often terminate the currently-executing task on these processors, resulting in wasted cycles. Fault-tolerant scheduling will be critical to the scalability of future HEC systems to an extremely large number of processors. The increasing usage of high-end systems for both computation- and data-intensive applications requires that future scheduling systems must address the problem of co-scheduling CPUs, I/O and network resources when mapping tasks to appropriate resources. In some HEC architectures, the placement of processes has an effect on the overall performance of the application due to the underlying interconnection topology. Both co-scheduling and workload-aware scheduling will be important for future high-end computing systems.An integrated software framework will be developed for scheduling and resource management for extreme-scale computing systems that provides the following capabilities: (i) on-line workload characterization, (ii) predictive scheduling based on time-series modeling and forecasting of resource utilization and queued workload, (iii) transparent fault-tolerant scheduling of applications that are interrupted by one or more faults, and (iv) efficient heuristic and evolutionary algorithms that consider the workload and resource characteristics and forecasts as part of the scheduling decision.This work represents a mix of scheduling theory and robust software development. These algorithms and the software framework in a production HEC environment at the Center for Advanced Computing (CAC) at the University of Michigan. The CAC HEC facility provides a testbed consisting of over 1400 CPUs representingmultiple processor families (AMD Athlons, Opterons, Apple Xserve/G5) and several interconnection systems such as Gigabit Ethernet and Myrinet.The intellectual merit of this research will be to advance the state-of-the-art in scheduling and resource management for HEC systems. By implementing and deploying the proposed framework at CAC, we will be able to collect realistic workload traces from a diverse array of end-users and applications. This research will serve as a catalyst for the development of robust fault-tolerant scheduling algorithms and their implementation software for such systems. We specifically address the scalability of fault-tolerance mechanisms such as checkpointing/restartand I/O that can scale across thousands of processors. Furthermore, our proposed integration of research, outreach and education activities will make broader impacts to other HEC centers and the scheduling research community. The framework developed as part of this project and our research results will be disseminated to industry and academia through open-source software and high-quality publications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: SaTC: CORE: Medium: Securing Interactions between Driver and Vehicle Using Batteries
CPS:Small: Imposing Recovery Period for Battery Health Monitoring, Prognosis, and Optimization
CPS: Breakthrough: Secure Interactions with Internet of Things
CPS: Synergy: Adaptive Management of Large Energy Storage Systems for Vehicle Electrification
国内基金
海外基金
Next Generation Majorana Nanowire Hybrids