Collaborative Research: ITR/NGS: Deja Vu: Transparent Checkpointing and Migration of Parallel Codes Over Grid Infrastructures
Collaborative Research: ITR/NGS: Deja Vu: Transparent Checkpointing and Migration of Parallel Codes Over Grid Infrastructures
批准号:
0325182
负责人:
Nathan Stone
金额:
$26.03万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-04-15 至 2009-03-31
中文摘要
从今天的计算网格到真正的网络基础设施的演变是一个艰巨的挑战,它无缝地集成了从学术实验室的小型集群到最大的国家超级计算中心的资源,并提供无处不在的高性能计算、研究仪器、数据仓库和可视化访问。实现这一未来需要在透明故障恢复机制方面取得根本性进展,以掩盖任何大规模计算资源特有的组件故障。虽然前几代超级计算机将可靠性设计到系统硬件中,但今天的高性能计算(HPC)环境是基于COTS组件集群的,对于整个资源的可靠性没有系统的解决方案。在不断增长的集群系统网络集合中产生稳定性需要一个软件解决方案,该解决方案通过透明、高效和自动的检查点和恢复(CPR)机制提供对计算资源的可靠访问。该项目旨在通过全新的方法来解决长期存在的CPR问题,并通过构建一个称为Deja vu的集成系统来实现迁移过程。Deja vu提供了(a)透明的并行检查点和恢复机制,可以从任何系统故障组合中恢复,而无需对并行应用程序进行任何修改。(b)一种新的编译后分析系统,透明地捕获应用程序状态;(c)一种系统架构,无缝地将用户发起和系统发起的检查点集成在一个框架中,从而有效地使用各种领域特定知识;(d)一种新的运行时机制,用于透明的增量检查点,以有效地捕获保持全局一致性所需的最少状态;(e)一种新颖的通信架构,可以透明地迁移现有的MPI/PVM代码,而无需修改应用程序或MPI/PVM库的源代码;(f)可针对特定存储环境进行定制的可恢复IO子系统;(g)与Globus Toolkit的接口和增强,以有效地使用本研究提供的CPR和迁移功能。“似曾相识”系统的核心CPR和迁移设施将被管理、安全和调度设施所包围,这些设施(a)与本地调度系统(例如OpenPBS)和会计系统集成,用于特定地点的会计和对损失的计算周期进行补偿;(b)扩展Globus安全架构,提供细粒度权限和动态创建的用户帐户,使“似曾相识”系统下可用的流体资源控制得到充分利用。这个项目的设计目标不仅仅是实现“点”解决方案,而是一个集成系统,它将构成大规模计算设施和网格基础设施的基本组件。我们的研究团队(VT, PSC, ISR)在完整解决方案的设计,开发,部署和支持方面具有丰富的经验。
英文摘要
A daunting challenge is the evolution from today's computational Grid to a true cyberinfrastructure that seamlessly integrates resources ranging from small clusters in academic laboratories to the largest national supercomputing centers and provides ubiquitous access to high performance computing, research instrumentation, data warehouses and visualization. Realization of this future requires fundamental advances in transparent fault recovery mechanisms to mask component failures endemic to any large-scale computational resource. While previous generations of supercomputers engineered reliability into systems hardware, today's high performance computing (HPC) environments are based on clusters of COTS components, with no systemic solution for the reliability of the resource as a whole. Engendering stability in ever growing networked collections of cluster systems needs a software solution that provides reliable access to computing resources through transparent, efficient, and automatic checkpointing and recovery (CPR) mechanisms. This project aims to bring about this future through radically new approaches to longstanding problems in CPR and process migration by building an integrated system called Deja vu. Deja vu provides (a) a transparent parallel checkpointing and recovery mechanism that recovers from any combination of systems failures without any modification to parallel applications. (b) a novel post-compiler analysis system that transparently captures application state, (c) a systems architecture that seamlessly integrates user-initiated and system-initiated checkpoints in a single framework enabling the effective use of a wide variety of domain specific knowledge, (d) novel runtime mechanisms for transparent incremental checkpointing, to efficiently capture the least amount of state required to maintain global consistency, (e) a novel communications architecture that enables transparent migration of existing MPI/PVM codes without source-code modifications to either the application or the MPI/PVM libraries, (f) recoverable IO subsystems that can be tailored to specific storage environments, and (g) interfaces to and augmentation of the Globus Toolkit to effectively use the CPR and migration capabilities provided by this research. The core CPR and migration facilities of Deja vu will be surrounded by management, security, and scheduling facilities that (a) integrate with local scheduling systems (e.g., OpenPBS) and accounting systems for site-specific accounting and refunding of lost compute cycles and (b) extend the Globus security architecture with fine grain rights and dynamically created user accounts that allow the fluid resource control available under the Deja vu system to be fully exploited. The design goal of this project is not just to implement "point" solutions, but an integrated system that will constitute a fundamental component of both large-scale computing facilities and Grid infrastructures. Our research team (VT, PSC, ISR) has considerable experience in the design, development, deployment and support of complete solutions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: