CSR-PSCE,SM: Recovery Aware Parallel Computing
CSR-PSCE,SM: Recovery Aware Parallel Computing
批准号:
0834514
负责人:
Zhiling Lan
金额:
$33.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-01 至 2013-05-31
中文摘要
随着并行系统的规模和复杂性不断增长,故障是不可避免的。多年来的研究主要集中在故障预预测和容限-预测故障并在故障发生前采取预防措施。尽管在故障预测方面取得了进展,但在实践中,特别是在具有前所未有的规模和复杂性的现代系统中,仍会发生意想不到的故障。由于故障的必然性,仅依靠故障前预测和容错是不足以进行故障管理的。正如故障发生时需要小心地避免和管理一样,故障后诊断和恢复同样重要,并且对并行计算的几乎每个方面都有深远的影响。本研究项目的目标是开发RAPS,一种用于故障后诊断和恢复的恢复感知并行计算系统。研究的重点是如何在发生故障后快速有效地恢复并行计算。最终目标是将故障后诊断和恢复与故障前预测和容错无缝集成,作为并行计算的复合故障管理解决方案。该方法包括(1)开发用于快速故障检测和根本原因分析的新诊断机制,(2)开发用于恢复协调的全系统编排,(3)设计用于并行应用快速恢复的新恢复技术,以及(4)综合评估。该项目的研究结果可以显著提高并行系统的生产率。这个项目也加强了印度理工学院的计算机科学课程,并扩大了代表性不足群体的参与。
英文摘要
As the scale and complexity of parallel systems continue to grow, failures are inevitable. For years research focused on pre-failure prediction and tolerance - predicting failures and taking precautionary actions before failure occurrence. Despite progress on failure prediction, unexpected failures occur in practice, especially in modern systems with unprecedented sizes and complexities. Relying on pre-failure prediction and tolerance alone is insufficient for fault management because of the inevitability of failures. Just as failures need to be carefully avoided and managed when they occur, post-failure diagnosis and recovery is of equal importance and has a profound impact on almost every aspect of parallel computing. The goal of this research project is to develop RAPS, a Recovery Aware Parallel computing System for post-failure diagnosis and recovery. The research focuses on how to quickly and effectively resume parallel computing after a failure has occurred. The ultimate goal is to seamlessly integrate post-failure diagnosis and recovery with pre-failure prediction and tolerance as a compound fault management solution for parallel computing. The approach consists of (1) development of new diagnosis mechanisms for fast failure detection and root cause analysis, (2) development of system-wide orchestration for recovery coordination, (3) design of new recovery techniques for quick restoration of parallel applications, and (4) a comprehensive evaluation. The results of this project can significantly improve the productivity of parallel systems. This project also enhances the CS curriculum at IIT and broadens the participation by underrepresented groups.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SHF:Small:Intelligent Management of Hybrid Workloads for Extreme Scale Computing
-
批准号:2413597
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2023
-
负责人:Zhiling Lan
-
依托单位:
Collaborative Research: PPoSS: Planning: SEEr: A Scalable, Energy Efficient HPC Environment for AI-Enabled Science
-
批准号:2119294
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2021
-
负责人:Zhiling Lan
-
依托单位:
SHF:Small:Intelligent Management of Hybrid Workloads for Extreme Scale Computing
-
批准号:2109316
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2021
-
负责人:Zhiling Lan
-
依托单位:
CSR: Small: IRON: Reducing Workload Interference on Massively Parallel Platforms
-
批准号:1717763
-
项目类别:Standard Grant
-
资助金额:$49.72万
-
财政年份:2017
-
负责人:Zhiling Lan
-
依托单位:
SHF: Small: Collaborative Research: Experimental-based Research on Effective Models of Parallel Application Execution Time, Power, and Resilience
-
批准号:1618776
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2016
-
负责人:Zhiling Lan
-
依托单位:
SHF: CSR: Small: Toward Smart HPC through Active Learning and Intelligent Scheduling
-
批准号:1422009
-
项目类别:Standard Grant
-
资助金额:$49.88万
-
财政年份:2014
-
负责人:Zhiling Lan
-
依托单位:
SHF: CSR: Small: A Cooperative Framework for Topology Awareness on Large-Scale Systems
-
批准号:1320125
-
项目类别:Standard Grant
-
资助金额:$49.84万
-
财政年份:2013
-
负责人:Zhiling Lan
-
依托单位:
Collaborative Research: Towards Petascale Cosmological Simulations
-
批准号:0904670
-
项目类别:Standard Grant
-
资助金额:$34.58万
-
财政年份:2009
-
负责人:Zhiling Lan
-
依托单位:
CSR/AES: Enhancing Application Robustness via Adaptive and Cooperative Methods
-
批准号:0720549
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Zhiling Lan
-
依托单位:
海外基金