课题基金 / 基金详情

CSR-DMSS, PSCE: Collaborative Research: Scalable Resilience in Large-Scale Systems

CSR-DMSS, PSCE: Collaborative Research: Scalable Resilience in Large-Scale Systems
CSR-DMSS、PSCE:协作研究:大型系统中的可扩展弹性
批准号:
0834483
负责人:
Chokchai Leangsuksun
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-01 至 2012-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目的目标是系统地设计从极端规模的高性能计算(HPC)分布中获取、标准化和处理量化的可靠性、可用性和可维护性(RAS)信息的方法,并通过研究和创建覆盖整个计算环境的最优反馈控制回路,为这些系统的实时RAS监控和建模开发一个新的、可扩展的框架。这项工作是由于HPC系统的规模和范围的持续和实质性的增加,这导致这些机器遇到的故障、错误和其他性能中断的数量迅速膨胀。随着HPC系统迈向Petaflop时代,必须更多地关注这些机器遇到的性能中断,以及它们可以继续不间断计算的方法的发展。在这种极端规模的环境中,旨在保持高可靠性和正常运行时间的努力是徒劳的吗?由于它们的处理器和计算单元数量巨大,这些系统不可避免地会遇到性能问题,而且失败势在必行。该项目旨在1)研究和开发先进的标准化方法,用于收集应用程序和系统级别的数据并生成可量化的RAS指标;2)提供一种新颖的、可扩展的解决方案,以提高在可靠预测大规模系统中迫在眉睫的节点和系统故障方面的准确性;以及3)设计防御性和主动性技术,以降低及时和准确处理弹性问题和对系统健康进行建模所需的计算成本。总之,这项工作试图缓解当代反应式容错方案的时间和成本限制,并将推动大规模计算部署中可扩展、主动和智能弹性供应的发展。此外,还将成立复原力联盟,以协同研究和开发,分享数据和结果,并向公众传播知识。
英文摘要
The objective of this project is to systematically design means of obtaining, tandardizing, and manipulating quantified Reliability, Availability and Serviceability (RAS) information from extreme-scale High Performance Computing (HPC) distributions, and to develop a novel, scalable framework for the real-time RAS monitoring and modeling of these systems via the research and creation of an optimal feedback control loop encompassing the entire computational environment. This work is necessitated by the continual and substantial increase in the size and scope of HPC systems, which is causing rapid inflation in the number of faults, errors, and other performance interruptions encountered by these machines.As HPC systems move towards the petaflop era, a greater focus must be placed on the performance interruptions encountered by these machines, and the development of means by which they may continue uninterrupted computation. In this extreme-scale environment, efforts aimed towards maintaining high reliability and uptime are futile ? with their enormous processor and computational unit counts, these systems will inevitably encounter performance issues, and failure must be expected. This project aims to 1) research and develop advanced, standardized methodologies for gathering application- and system-level data and generating quantifiable RAS metrics, 2) provide a novel, scalable solution for improving accuracy in reliably predicting imminent node-wise and system failures in large-scale systems, and 3) devise defensive and proactive techniques for reducing the computational costs required to timely and accurately handle resilience issues and model system health. In summary, this work attempts to alleviate the time and cost limitations of contemporary, reactive fault tolerance schemes, and will advance the development of scalable, proactive, and intelligent resilience provision in large-scale computing deployments. In addition, the Resilience Consortium will be established to synergistically research and develop, share data and findings, and disseminate knowledge to the public.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金