CSR-DMSS, PSCE: Collaborative Research: Scalable Resilience in Large-Scale Systems
CSR-DMSS, PSCE: Collaborative Research: Scalable Resilience in Large-Scale Systems
批准号:
0834483
负责人:
Chokchai Leangsuksun
金额:
$30.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-01 至 2012-08-31
中文摘要
该项目的目标是系统地设计从超大规模高性能计算(HPC)分布中获取、标准化和操纵量化的可靠性、可用性和可服务性(RAS)信息的方法,并通过研究和创建包含整个计算环境的最佳反馈控制回路,为这些系统的实时RAS监控和建模开发一种新颖的、可扩展的框架。由于高性能计算系统的规模和范围不断大幅增加,这项工作是必要的,这导致这些机器遇到的故障、错误和其他性能中断的数量迅速增加。随着HPC系统走向千万亿次的时代,必须更加关注这些机器所遇到的性能中断,以及开发它们可以继续不间断计算的方法。在这种极端规模的环境中,旨在保持高可靠性和正常运行时间的努力是徒劳的?由于这些系统拥有庞大的处理器和计算单元数量,它们将不可避免地遇到性能问题,并且必然会出现故障。该项目旨在1)研究和开发先进的、标准化的方法,用于收集应用程序和系统级数据,并生成可量化的RAS指标;2)提供一种新颖的、可扩展的解决方案,以提高在大规模系统中可靠地预测即将发生的节点和系统故障的准确性;3)设计防御和主动技术,以减少及时、准确地处理弹性问题和系统健康模型所需的计算成本。总之,这项工作试图减轻当代反应性容错方案的时间和成本限制,并将推动大规模计算部署中可扩展、主动和智能弹性供应的发展。此外,还将建立复原力联盟,以协同研究和开发,共享数据和发现,并向公众传播知识。
英文摘要
The objective of this project is to systematically design means of obtaining, tandardizing, and manipulating quantified Reliability, Availability and Serviceability (RAS) information from extreme-scale High Performance Computing (HPC) distributions, and to develop a novel, scalable framework for the real-time RAS monitoring and modeling of these systems via the research and creation of an optimal feedback control loop encompassing the entire computational environment. This work is necessitated by the continual and substantial increase in the size and scope of HPC systems, which is causing rapid inflation in the number of faults, errors, and other performance interruptions encountered by these machines.As HPC systems move towards the petaflop era, a greater focus must be placed on the performance interruptions encountered by these machines, and the development of means by which they may continue uninterrupted computation. In this extreme-scale environment, efforts aimed towards maintaining high reliability and uptime are futile ? with their enormous processor and computational unit counts, these systems will inevitably encounter performance issues, and failure must be expected. This project aims to 1) research and develop advanced, standardized methodologies for gathering application- and system-level data and generating quantifiable RAS metrics, 2) provide a novel, scalable solution for improving accuracy in reliably predicting imminent node-wise and system failures in large-scale systems, and 3) devise defensive and proactive techniques for reducing the computational costs required to timely and accurately handle resilience issues and model system health. In summary, this work attempts to alleviate the time and cost limitations of contemporary, reactive fault tolerance schemes, and will advance the development of scalable, proactive, and intelligent resilience provision in large-scale computing deployments. In addition, the Resilience Consortium will be established to synergistically research and develop, share data and findings, and disseminate knowledge to the public.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金