Adaptive and Reliable ParallelComputing9 Networks of Workstations

Adaptive and Reliable ParallelComputing9 Networks of Workstations
复制标题

自适应且可靠的并行计算9 工作站网络

DOI:
--
复制
发表时间:
1997
期刊:
影响因子:
--
通讯作者:
R. Blumofe
R. Blumofe
中科院分区:
--
文献类型:
--
作者:
P. Lisiecki;R. Blumofe

文献摘要

被引文献

相似文献

本文介绍了一个运行时系统Cilk-Now的设计,该系统能在一个由多个Unix工作站组成的网络上自适应地、可靠地并行执行函数式Cilk程序。Cilk(发音为“Silk”)是C语言的并行多线程扩展,所有的Cilk运行时系统都采用了被证明有效的线程调度算法。Cilk-Now就是这样一个运行时系统,此外,Cilk-Now自动为Cilk程序的一个功能子集提供自适应和可靠的执行。通过自适应执行,我们的意思是每个Cilk程序动态地利用一组变化的否则空闲的工作站。通过可靠执行,我们的意思是Cilk-Now系统作为一个整体和每个执行的Cilk程序都能够容忍机器和网络故障。Cilk-现在提供了这些功能,同时程序保持故障无感知,这意味着Cilk程序员不需要为容错编写代码。在本文中,我们将重点介绍端到端的设计决策,并展示这些决策如何允许设计利用Cilk编程模型的高级算法属性来简化和流线化实现。
In this paper, we present the design of Cilk-NOW, a runtime system that adaptively and reliably executes functional Cilk programs in parallel on a network of UNIX workstations. Cilk (pronounced "silk") is a parallel multithreaded extension of the C language, and all Cilk runtime systems employ a provably efficient threadscheduling algorithm. Cilk-NOW is such a runtime system, and in addition, Cilk-NOW automatically delivers adaptive and reliable execution for a functional subset of Cilk programs. By adaptive execution, we mean that each Cilk program dynamically utilizes a changing set of otherwise-idle workstations. By reliable execution, we mean that the Cilk-NOW system as a whole and each executing Cilk program are able to tolerate machine and network faults. Cilk-NOW provides these features while programs remain fault oblivious, meaning that Cilk programmers need not code for fault tolerance. Throughout this paper, we focus on end-to-end design decisions, and we show how these decisions allow the design to exploit high-level algorithmic properties of the Cilk programming model in order to simplify and streamline the implementation.