课题基金 / 基金详情

Tolerating faults in interconnection networks for parallel computing

Tolerating faults in interconnection networks for parallel computing
并行计算互连网络中的容错
批准号:
EP/G010587/1
负责人:
Iain Stewart
金额:
$35.07万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2009
资助国家:
英国
项目状态:
已结题
起止时间:
2009 至 --

项目摘要

项目成果

Iain Stewart的其他基金

相似基金

相关文献

中文摘要
翻译
在分布式内存多处理器中,抑制更快全局计算的主要因素是处理器间通信。通信依赖于互连网络的拓扑结构、路由机制、流量控制策略和交换方法。我们关心的是与互连网络的拓扑结构有关的问题。选择如何在分布式内存多处理器中连接处理器是一个基本的设计决策。有许多,往往是相互冲突的,要记住的考虑因素。例如,我们希望我们的互连网络是对称的(使编程和分析更容易),具有小直径(以减少消息传递延迟),可递归分解(以帮助可扩展性),高度连接(以提高容错性和可靠性),是低程度的规则(以减少通信开销和设计复杂性),支持快速和容易的处理器间通信,支持基于其他拓扑的其他机器的仿真。等等......。这些特性都提高了计算性能。然而,不存在一个在所有方面都是最优的互连网络,必须做出权衡。许多互连网络已经被提出,每一个网络都有一些好的(拓扑)性质和一些不太好的。当构建具有大量处理器的分布式内存多处理器时,需要一定的容错能力,因为人们仍然希望机器在(有限数量的)处理器或链路故障下运行。至于在容错方面需要什么取决于上下文,但最低要求是互连网络(非故障部分)应保持连接。然而,通常需要更多。与并行计算相关的其他重要性质包括哈密顿性质,因为网络中哈密顿循环的存在是至关重要的,因为在许多分布式算法中,这种循环作为数据结构无处不在(它们主要用于促进消息传递)。不仅哈密顿环的存在很重要,哈密顿路径的存在也很重要,更一般地说,不同长度的环和路径的存在也很重要。哈密顿路径(或者至少是长路径)的存在是非常有用的,因为我们经常需要在分布式内存多处理器中模拟线性数组计算;拥有一条较长的路径可以让我们迎合模拟中包含许多不同数组长度的模拟。此外,考虑到并行计算中基于循环的计算和算法的普遍存在,不仅基于线性数组的计算的模拟很重要,而且(不同长度的)基于循环的计算的模拟也很重要。本文主要研究的是互连网络中的故障容错问题。该研究有三个主线:在条件故障假设(即对网络中故障分布的假设)下,研究各种互连网络中路径和(不同长度的)循环的存在性;光转置互联系统(OTIS)网络容错性研究以及在故障网络中分布式构建嵌入式结构。
英文摘要
In distributed-memory multiprocessors, the dominant factor inhibiting faster global computations is inter-processor communication. Communication is dependent upon the topology of the interconnection network, the routing mechanism, the flow control policy , and the method of switching. We are concerned with issues relating to the topology of the interconnection network. The choice of how we connect processors in a distributed-memory multiprocessor is a fundamental design decision. There are numerous, often conflicting, considerations to bear in mind. For instance, we would like our interconnection network to be symmetric (to make programming and analysis easier), have small diameter (to lessen message-passing latency), be recursively decomposable (to aid scalability), be highly connected (to improve fault-tolerance and reliability), be regular of low degree (to lessen communication overheads and design complexity), support rapid and easy inter-processor communication, support the simulation of other machines based on other topologies, and so on. These properties all give rise to improved computational performance. However, there does not exist an interconnection network that is optimal on all counts and trade-offs have to be made. A multitude of interconnection networks have been proposed with each of these networks having some good (topological) properties and some not so good. When building distributed-memory multiprocessors with massive numbers of processors, some capacity for fault-tolerance is required, for one would still wish the machine to be operative under (a limited number of) processor or link faults. As to what one requires in terms of fault-tolerance depends upon the context, but the minimal requirement is that the (non-faulty portion of the) interconnection network should remain connected. However, usually more is required. Other important properties relevant to parallel computing include Hamiltonicity properties, for the existence of Hamiltonian cycles in networks is of crucial importance, given the ubiquity of such cycles as data structures in many distributed algorithms (they are primarily used to facilitate message-passing). Not only is the existence of Hamiltonian cycles of great importance but also the existence of Hamiltonian paths, and more generally the existence of cycles and paths of different lengths. The existence of Hamiltonian (or, at least, long) paths is extremely useful as we regularly need to simulate linear-array computations in distributed-memory multiprocessors; having a long path allows us to cater for such simulations where there are many different array-lengths involved in the simulations. In addition, given the ubiquity of cycle-based computations and algorithms in parallel computation, not only is the simulation of linear-array-based computations important but so is the simulation of cycle-based computations (of varying lengths).The research in this proposal is all about the toleration of faults in interconnection networks. There are three threads to the research: the study of the existence of paths and cycles (of varying lengths) in various interconnection networks under conditional fault assumptions (that is, asumptions on the distributions of the faults in the network); the study of fault-tolerance in Optical Transpose Interconnect System (OTIS) networks; and the distributed construction of embedded structures within a faulty network.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1016/j.tcs.2012.05.015
发表时间: 2012-09
期刊: Theor. Comput. Sci.
影响因子: --
作者: [I. A. Stewart]
通讯作者: I. A. Stewart
DOI: 10.1142/s012962641250003x
发表时间: 2012-04
期刊: Parallel Process. Lett.
影响因子: --
作者: [I. A. Stewart]
通讯作者: I. A. Stewart
ALGOUK - A Network for Algorithms and Complexity in the UK
  • 批准号:
    EP/R005613/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $13.85万
  • 财政年份:
    2017
  • 负责人:
    Iain Stewart
  • 依托单位:
Interconnection Networks: Practice unites with Theory (INPUT)
  • 批准号:
    EP/K015680/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $45.05万
  • 财政年份:
    2013
  • 负责人:
    Iain Stewart
  • 依托单位:
Quantified Constraints and Generalisations
  • 批准号:
    EP/G020604/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $31.54万
  • 财政年份:
    2009
  • 负责人:
    Iain Stewart
  • 依托单位:
Finite and Algorithmic Model Theory
  • 批准号:
    EP/D056853/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $2.42万
  • 财政年份:
    2006
  • 负责人:
    Iain Stewart
  • 依托单位:
国内基金
海外基金
制冷系统故障诊断关键问题的定量研究
  • 批准号:
    50876059
  • 项目类别:
    面上项目
  • 资助金额:
    30.0万元
  • 批准年份:
    2008
  • 负责人:
    谷波
  • 依托单位: