Reliability, Performability and Scalability of Large-Scale Distributed Systems
Reliability, Performability and Scalability of Large-Scale Distributed Systems
批准号:
9010240
负责人:
Walid Najjar
金额:
$6.42万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1990
资助国家:
美国
项目状态:
已结题
起止时间:
1990-07-01 至 1992-12-31
中文摘要
大规模分布式多计算机系统,由多个 数以千计的处理元件,正在迅速证明他们的 作为低成本高性能超级计算机的潜力。 不仅可以 这些系统加速了程序的执行,但它们也允许 更大的问题需要解决。 广泛使用 然而,这些系统在使命关键领域和商业领域都是如此 应用程序,取决于其已证明的可靠性、可用性 和可扩展性。 本研究项目的目的是调查 大规模分布式系统的可靠性、可扩展性和可执行性 系统. 随着系统中元素数量的增加, 系统的故障预计会增加, 技术. 因此,系统的可靠性和可扩展性 在大型系统设计中的重要考虑。 的 研究将集中在两个基本问题:网络分析 可靠性和可执行性,以及评估技术, 可以利用这些系统固有的冗余。 网络 可靠性分析将检查多个节点的影响, 链路故障对网络的连通性及其 通信带宽,调查发生概率 网络断开、饱和和通信瓶颈。 大规模系统固有的硬件冗余可以 利用,以实现更高的可靠性,虽然,在成本 降低计算能力。 第二个目标是调查 可实现性能/可靠性折衷和系统可扩展性 使用各种冗余方案。 这项研究基本上是分析性质的,但将依赖于 模拟技术,只要一个精确的分析评估是不 可行
英文摘要
Large-scale distributed multicomputer systems, consisting of several thousand processing elements, are rapidly demonstrating their potential as a low cost high performance supercomputer. Not only can these system speed-up program execution, but they also allow significantly larger problems to be addressed. A wide-spread use of these systems, however, in mission critical as well as commercial applications, depends on their demonstrated reliability, availability and scalability. The objective of this research project is to investigate the reliability, scalability and performability of large-scale distributed systems. As the number of elements in a system increases, the rate of failure of the system is expected to increase given a constant technology. Therefore system reliability and scalability are important considerations in the design of large-scale systems. The research will focus on two essential issues: the analysis of network reliability and performability, and the evaluation of techniques that can exploit the inherent redundancy of these systems. The network reliability analysis will examine the effects of multiple node and link failures on the connectivity of the network and on its communication bandwidth, investigating the probability of occurrence of network disconnection, saturation and communication bottlenecks. The inherent hardware redundancy of large-scale systems can be exploited to achieve a higher reliability, albeit, at the cost of a reduced computing power. The second objective will be to investigate the achievable performance/reliability tradeoff and system scalability using various redundancy schemes. The research is essentially analytical in nature but will rely on simulation techniques whenever an exact analytical evaluation is not feasible.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SHF:Small: Automatic Generation of Hardware Threads on Programmable Fabrics
-
批准号:1219180
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2012
-
负责人:Walid Najjar
-
依托单位:
CPA-CSA: Hardware Support for FPGA-Based Code Acceleration
-
批准号:0811416
-
项目类别:Continuing Grant
-
资助金额:$23.0万
-
财政年份:2008
-
负责人:Walid Najjar
-
依托单位:
SGER: Hardward/Software Partitioning for Multiprocessor and Multicore Acceleration
-
批准号:0745490
-
项目类别:Standard Grant
-
资助金额:$3.69万
-
财政年份:2007
-
负责人:Walid Najjar
-
依托单位:
海外基金