课题基金 / 基金详情

SPX: Collaborative Research: Cross-layer Application-Aware Resilience at Extreme Scale (CAARES)

SPX: Collaborative Research: Cross-layer Application-Aware Resilience at Extreme Scale (CAARES)
SPX:协作研究:超大规模跨层应用程序感知弹性 (CAARES)
批准号:
1725649
负责人:
Ivan Rodero
金额:
$26.72万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-15 至 2020-07-31

项目摘要

项目成果

Ivan Rodero的其他基金

相似基金

相关文献

中文摘要
翻译
科学和工程应用日益增长的需求推动了当前大规模系统的极限,预计在下一个十年的早期将实现exascale(10^18 FLOPS)性能。在极端规模下较少研究的挑战之一是计算系统本身的可靠性,主要是由于使用了非常大量的核心和组件,以及这些系统上的平均故障间隔时间急剧减少(大约几十分钟)。该项目从传统的单组件故障管理模型出发,并探讨如何在一个单一的并行应用程序的上下文中使用的多个软件库(和应用程序组件)可以交互,以提供必要的并行应用程序的能力计算的整体故障管理支持。这种探索将不限于使用单一的并行编程范式开发的软件,但将扩展到包括更具挑战性的情况下,多个编程范式可以结合起来,以实现一个共同的目标,模拟一组大规模的科学应用程序在今天使用。 该项目的目标是摆脱当前孤立的弹性机制,并提出跨层组合解决方案,从根本上解决这些极端规模的弹性挑战。这种探索将不限于使用单一的并行编程范式开发的软件,但将扩展到包括更具挑战性的情况下,多个编程范式可以结合起来,以实现一个共同的目标,模拟一组大规模的科学应用程序在今天使用。更具体地说,这项建议将解决以下研究挑战:(1)为更深入地了解不同复原力模型和方法相结合所带来的挑战和机遇奠定理论基础;(2)设计灵活的方案编制抽象概念,以便将不同的复原力模型和机制相结合,以更全面的方式开展合作和解决复原力问题;以及(3)开发基本的、独立于编程范式的、实现跨层和特定领域方法所必需的结构,以支持弹性并理解相关的性能/质量权衡。所提出的方法将通过在两种不同的编程范式(MPI和OpenSHMEM)中暴露这些通用抽象来验证,通过为这些范式中的每一个创建和开发专门的概念。这将使评估的有效性的概念和相应的间接费用所施加的不同的软件层,使用几个软件框架和应用程序。
英文摘要
The increasing demands of science and engineering applications push the limits of current large-scale systems, and is expected to achieve exascale (10^18 FLOPS) performance early in the next decade. One of the lesser studied challenge at extreme scales is the reliability of the computing system itself, primarily due to the very large number of cores and components utilized and to the sharp decrease of the Mean Time Between Failures on such systems (in the order of tens of minutes). This project departs from the traditional single component fault management model, and explores how multiple software libraries (and application components) used in the context of a single parallel application can interact to provide the holistic fault management support necessary for parallel applications targeting capability computing. This exploration will not be limited to software developed using a single parallel programming paradigm, but will be extended to encompass the more challenging case where multiple programming paradigms can be combined to achieve a common goal, to simulate a set of large scale scientific applications in use today.  The goal of this project is to depart from the current siloed resilience mechanisms, and propose cross-layer composition solutions that can fundamentally address these resilience challenges at extreme scales. This exploration will not be limited to software developed using a single parallel programming paradigm, but will be extended to encompass the more challenging case where multiple programming paradigms can be combined to achieve a common goal, to simulate a set of large scale scientific applications in use today. More specifically, this proposal will address the following research challenges: (1) development of a theoretical foundation for a deeper understanding of the challenges and opportunities arising from combining different resilience models and methodologies; (2) design of a flexible programming abstraction to allow different resilience models and mechanisms to be combined to cooperate and address resilience in a more holistic manner; and (3) development of basic, programming paradigm independent, constructs necessary to implement cross-layer and domain-specific approaches to support resilience and to understand related performance / quality trade-offs. The proposed approach will be validated by exposing these generic abstractions in two different programming paradigms (MPI and OpenSHMEM), by creating and developing specialized concepts for each of these paradigms. This will enable the assessment of the validity of the concepts and the corresponding overheads imposed by the different software layers, using few software frameworks and applications.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Scalable Crash Consistency for Staging-based In-situ Scientific Workflows
基于分期的原位科学工作流程的可扩展崩溃一致性
DOI: 10.1109/ipdpsw50202.2020.00068
发表时间: 2020
期刊: 2020 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW
影响因子: --
作者: [Duan, Shaohua, Parashar, Manish]
通讯作者: Parashar, Manish
DOI: 10.1109/ipdps.2018.00021
发表时间: 2018-05
期刊: 2018 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子: --
作者: [Shaohua Duan;P. Subedi;K. Teranishi;Philip E. Davis;H. Kolla;Marc Gamell;M. Parashar]
通讯作者: Shaohua Duan;P. Subedi;K. Teranishi;Philip E. Davis;H. Kolla;Marc Gamell;M. Parashar
CIF21 DIBBs: EI: Virtual Data Collaboratory: A Regional Cyberinfrastructure for Collaborative Data Intensive Science
  • 批准号:
    2220826
  • 项目类别:
    Standard Grant
  • 资助金额:
    $400.0万
  • 财政年份:
    2021
  • 负责人:
    Ivan Rodero
  • 依托单位:
Collaborative Research: Framework: Data: NSCI: HDR: GeoSCIFramework: Scalable Real-Time Streaming Analytics and Machine Learning for Geoscience and Hazards Research
  • 批准号:
    2219975
  • 项目类别:
    Standard Grant
  • 资助金额:
    $89.91万
  • 财政年份:
    2021
  • 负责人:
    Ivan Rodero
  • 依托单位:
Collaborative Research: Framework: Data: NSCI: HDR: GeoSCIFramework: Scalable Real-Time Streaming Analytics and Machine Learning for Geoscience and Hazards Research
  • 批准号:
    1835692
  • 项目类别:
    Standard Grant
  • 资助金额:
    $89.91万
  • 财政年份:
    2019
  • 负责人:
    Ivan Rodero
  • 依托单位:
NSF Large Facilities Cyberinfrastructure Workshop
  • 批准号:
    1742969
  • 项目类别:
    Standard Grant
  • 资助金额:
    $6.51万
  • 财政年份:
    2017
  • 负责人:
    Ivan Rodero
  • 依托单位:
海外基金