Cumulvs: Providing Fault Toler. Ance, Visualization, and Steer Ing of Parallel Applications

Cumulvs: Providing Fault Toler. Ance, Visualization, and Steer Ing of Parallel Applications
复制标题

Cumulvs:提供容错。

DOI:
10.1177/109434209701100305
复制
发表时间:
1996
影响因子:
3.1
通讯作者:
P. Papadopoulos
P. Papadopoulos
中科院分区:
计算机科学3区
文献类型:
--
作者:
A. Geist;J. Kohl;P. Papadopoulos

文献摘要

被引文献

相似文献

可视化和计算指导的使用通常可以帮助科学家分析大规模的科学应用。在分布式系统上运行时,对故障的容错非常重要。然而,实现这些功能的细节复杂且繁琐,导致许多科学家没有足够的开发工具。 CUMULVS 是一个库,使程序员能够轻松地将交互式可视化和计算引导合并到现有的并行程序中。 CUMULVS 基于 PVM 虚拟机框架构建,具有可移植性,并且可与 PVM 所使用的所有计算机体系结构互操作——这个列表还在不断增加,目前约有 60 种体系结构。 CUMULVS 库分为两部分:一部分用于应用程序,另一部分用于可能的商业、可视化和转向前端。这两个库共同传递了将多个独立查看器前端动态附加到正在运行的并行应用程序所需的所有连接和数据协议。查看器程序还可以引导一个或多个用户定义的参数来“闭合循环”以进行计算实验和分析。 CUMULVS 允许程序员指定用户控制的检查点,以便在出现故障时保存重要的程序状态,并且还提供了一种跨异构机器架构迁移任务的机制,以实现更高的性能。给出了 CUMULVS 设计目标和妥协以及未来方向的详细信息。
The use of visualization and computational steering can often assist scientists in analyzing large-scale scientific applications. Fault tolerance to failures is of great impor tance when running on a distributed system. However, the details of implementing these features are complex and tedious, leaving many scientists with inadequate develop ment tools. CUMULVS is a library that enables program mers to easily incorporate interactive visualization and computational steering into existing parallel programs. Built on the PVM virtual machine framework, CUMULVS is portable and interoperable with all the computer archi tectures that PVM works with—a growing list that now stands at about 60 architectures. The CUMULVS library is divided into two pieces: one for the application program and one for the possibly commercial, visualization, and steering front end. Together, these two libraries encom pass all the connection and data protocols needed to dynamically attach multiple, independent viewer front ends to a running parallel application. Viewer programs can also steer one or more user-defined parameters to "close the loop" for computational experiments and analy ses. CUMULVS allows the programmer to specify user- directed checkpoints for saving an important program state in case of failures and also provides a mechanism to migrate tasks across heterogeneous machine architec tures to achieve improved performance. Details of the CUMULVS design goals and compromises as well as future directions are given.