Graceful Termination -- Graceful Resetting

Graceful Termination -- Graceful Resetting
复制标题

优雅终止——优雅重置

DOI:
--
复制
发表时间:
1989
期刊:
影响因子:
--
通讯作者:
P. Welch
P. Welch
中科院分区:
--
文献类型:
--
作者:
P. Welch

文献摘要

被引文献

相似文献

正确的——更不用说优雅的——并行系统的终止有时被认为是一个难题。这在occam和transper网络的纯消息传递MIMD规则下尤其如此,其中不允许全局操作(如设置共享标志或流产),并且不能为每个通信设置超时。本文描述了解决该问题的一些常见但错误的occam方法,并将它们与Ada[0,1,2]中的方法进行了对比。由于不安全和性能开销,这些方法都被拒绝了。提出了一种简单、合法、安全、高效的occam方法。这种方法还解决了一个更重要的问题——并行系统(或子系统)的总体(或部分)重置。重置机制完全独立于并行应用算法,因此可以在开发时不必担心这些问题。这种关注点分离是良好的软件工程,并得到occam哲学的充分支持。最后,描述了该复位机制的一个应用,允许occam网络拓扑结构的动态重建。给定一个具有任意拓扑、消息协议和同步机制的进程(子)网络,安排其终止。主动终止系统可能来自一个或多个进程本身和/或来自外部的一个或多个点(如果网络不是封闭系统)。我们必须避免的陷阱是提交一个进程与已终止的邻居进行通信。如果发生这种情况,通信将永远不会终止,因此,网络将永远不会终止。如何避免——我给每个进程配备一个额外的中断通道。如果一个进程是作为子进程的并行网络来实现的,那么将这个中断通道“扇出”到每个子进程。如果一个进程有一个顺序实现,修改它的算法,以便当中断信号到达时终止。这一方案的问题在于,它的成功与否对关闭流程的顺序很敏感。具体来说,必须按照网络数据流的“拓扑”顺序终止进程。例如,如果网络是由四个进程组成的管道:−a、B、C、D中断。一个中断。interrupt意为中断。我们必须安排中断以A, B, C, D的顺序到达。假设interrupt.C在interrupt.B之前被触发。那么,很有可能进程C会在进程B注意到它的中断之前终止。在这种情况下,进程B可能会进行致命的输出消息尝试,并永远挂起。中断。B信号永远不会被确认。这可能会阻塞中断生成进程,从而使网络的许多其他部分仍然处于活动状态,从而造成进一步的损害。不幸的是,如果网络有反馈,就没有安全的顺序来触发这些中断——例如:−A B中断。一个中断。如果A在B之前终止,B可能会在尝试输出时卡住——反之亦然!在[1]中,A负责拉出中断。B(在它收到中断之后)。假定B检查它的中断。−•如果当A尝试中断时,B被承诺输出到A。B,通过让A总是中断来避免死锁。B与B的输入并行,A输出此中断。在B完成对a的正常输出之前,B将在B处挂起。这样,B下次将检测到它,而不会尝试再次与a通信。不幸的是,如果上述前提条件不为真,B将简单地检测它的中断。B和终止,不向A发送任何东西,留下A搁浅。当然,A可能会尝试对来自B的最后通信超时,但请参阅下一节。如何不这样做- 2决定杀死系统的进程只是终止。在所有最低级别(即顺序)网络进程上施加一些“超时”机制,以便它们在等待输入的阻塞时间足够长时放弃并死亡。该方案对正确设置超时值非常敏感。例如:−
Correct — let alone graceful — termination of parallel systems is sometimes thought to be a difficult problem. This is particularly imagined to be so under the pure message-passing MIMD discipline of occam and transputer networks, where global operations (like setting a shared flag or abortions) are not allowed and where time-outs cannot be set for every communication. This paper describes some common, but erroneous, occam approaches to this problem and contrasts them with what can be done in Ada [0, 1, 2]. These methods are all rejected on the grounds of insecurity and performance overheads. A simple, legal, secure and efficient occam method is then presented. This method also solves a much more important problem — the general (or partial) resetting of a parallel system (or sub-system). The resetting mechanism is quite independent of the parallel application algorithm, which can therefore be developed without worrying about such matters. This separation of concerns is good software engineering and is fully supported by the occam philosophy. Finally, an application of this resetting mechanism is described that permits the dynamic reconstruction of occam network topologies. The Problem Given a (sub-)network of processes with arbitrary topology, message protocol and synchronisation regime, arrange for it to terminate. The initiative to kill the system may come from one or more of the processes themselves and/or from one or more points outside (if the network is not a closed system). The pit-fall we have to avoid is committing a process to communicate with a terminated neighbour. If this were to happen, that communication would never terminate and, therefore, the network would never terminate. How Not To Do It — I Equip every process with an extra interrupt channel. If a process is implemented as a parallel network of sub-processes, ‘‘fan-out’’ this interrupt channel down to each of them. If a process has a sequential implementation, modify its algorithm so as to terminate if ever an interrupt signal arrives. The trouble with this scheme is that its success is sensitive to the order in which the processes are closed down. Specifically, the processes must be killed off in a ‘‘topological’’ ordering with respect to the network data-flow. For instance, if the network were a pipeline of four processes :− A B C D interrupt.A interrupt.B interrupt.C interrupt.D we must arrange for the interrupts to arrive in the sequence A, B, C, D. Suppose we did not. Suppose that interrupt.C fired before interrupt.B. Then, there is a good chance that process C will terminate before process B notices its interrupt.B. In that case, process B may make a fatal attempt to output a message and get suspended for ever. The interrupt.B signal never gets acknowledged. This might cause further damage by blocking the interrupt generating process — thus leaving many other parts of the network still active. Unfortunately, if the network has feed-back, there is no secure ordering possible for firing these interrupts — e.g. :− A B interrupt.A interrupt.B If A terminates before B, B may get stuck trying to output — and vice versa! In [1], A is given the responsibility for pulling the interrupt.B (after it has received an interrupt.A). B is assumed to check its interrupt.B line between every output to A. Then :− • If B were committed to output to A when A tries to interrupt.B, deadlock is avoided by having A always interrupt.B in parallel with inputting from B. A outputs this interrupt.B at high priority so that it will be pending at B before B completes its normal output to A. That way, B will detect it next time and not attempt to communicate again with A. • Unfortunately, if the above pre-condition were not true, B will simply detect its interrupt.B and terminate without ever sending anything back to A — leaving A stranded. Of course, A might try timing-out on this final communication from B — but see the next section. How Not To Do It — II The process that decides to kill the system off simply terminates. Impose some ‘‘time-out’’ mechanism on all the lowest level (i.e. sequential) network processes so that they simply give up and die if they are blocked long enough awaiting input. This scheme is very sensitive to setting the time-out values correctly. For instance, consider :−