Graceful Termination -- Graceful Resetting
Graceful Termination -- Graceful Resetting
复制标题
优雅终止——优雅重置
DOI:
--
复制
发表时间:
1989
期刊:
影响因子:
--
通讯作者:
P. Welch
中科院分区:
文献类型:
--
作者:
P. Welch
Correct — let alone graceful — termination of parallel systems is sometimes thought to be a difficult problem. This is particularly imagined to be so under the pure message-passing MIMD discipline of occam and transputer networks, where global operations (like setting a shared flag or abortions) are not allowed and where time-outs cannot be set for every communication. This paper describes some common, but erroneous, occam approaches to this problem and contrasts them with what can be done in Ada [0, 1, 2]. These methods are all rejected on the grounds of insecurity and performance overheads. A simple, legal, secure and efficient occam method is then presented. This method also solves a much more important problem — the general (or partial) resetting of a parallel system (or sub-system). The resetting mechanism is quite independent of the parallel application algorithm, which can therefore be developed without worrying about such matters. This separation of concerns is good software engineering and is fully supported by the occam philosophy. Finally, an application of this resetting mechanism is described that permits the dynamic reconstruction of occam network topologies. The Problem Given a (sub-)network of processes with arbitrary topology, message protocol and synchronisation regime, arrange for it to terminate. The initiative to kill the system may come from one or more of the processes themselves and/or from one or more points outside (if the network is not a closed system). The pit-fall we have to avoid is committing a process to communicate with a terminated neighbour. If this were to happen, that communication would never terminate and, therefore, the network would never terminate. How Not To Do It — I Equip every process with an extra interrupt channel. If a process is implemented as a parallel network of sub-processes, ‘‘fan-out’’ this interrupt channel down to each of them. If a process has a sequential implementation, modify its algorithm so as to terminate if ever an interrupt signal arrives. The trouble with this scheme is that its success is sensitive to the order in which the processes are closed down. Specifically, the processes must be killed off in a ‘‘topological’’ ordering with respect to the network data-flow. For instance, if the network were a pipeline of four processes :− A B C D interrupt.A interrupt.B interrupt.C interrupt.D we must arrange for the interrupts to arrive in the sequence A, B, C, D. Suppose we did not. Suppose that interrupt.C fired before interrupt.B. Then, there is a good chance that process C will terminate before process B notices its interrupt.B. In that case, process B may make a fatal attempt to output a message and get suspended for ever. The interrupt.B signal never gets acknowledged. This might cause further damage by blocking the interrupt generating process — thus leaving many other parts of the network still active. Unfortunately, if the network has feed-back, there is no secure ordering possible for firing these interrupts — e.g. :− A B interrupt.A interrupt.B If A terminates before B, B may get stuck trying to output — and vice versa! In [1], A is given the responsibility for pulling the interrupt.B (after it has received an interrupt.A). B is assumed to check its interrupt.B line between every output to A. Then :− • If B were committed to output to A when A tries to interrupt.B, deadlock is avoided by having A always interrupt.B in parallel with inputting from B. A outputs this interrupt.B at high priority so that it will be pending at B before B completes its normal output to A. That way, B will detect it next time and not attempt to communicate again with A. • Unfortunately, if the above pre-condition were not true, B will simply detect its interrupt.B and terminate without ever sending anything back to A — leaving A stranded. Of course, A might try timing-out on this final communication from B — but see the next section. How Not To Do It — II The process that decides to kill the system off simply terminates. Impose some ‘‘time-out’’ mechanism on all the lowest level (i.e. sequential) network processes so that they simply give up and die if they are blocked long enough awaiting input. This scheme is very sensitive to setting the time-out values correctly. For instance, consider :−