Robust Parallel and Distributed Computing Systems
Robust Parallel and Distributed Computing Systems
批准号:
0615170
负责人:
Howard Siegel
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-06-15 至 2010-05-31
中文摘要
并行和分布式计算系统由一组异构的机器、软件和网络组成,经常在由于不可预测的情况(例如突然的机器故障或系统参数估计的不准确)而导致其性能下降的环境中操作。 一个重要的问题是:偏离假设环境的程度会导致性能下降到系统无法满足指定要求的程度,即,该系统的稳健程度如何? 这项工作的重点是设计的方法,产生鲁棒性指标,并使用它们在资源管理。当特定系统参数发生某些扰动时,如果系统性能的退化保持在特定限度内,则资源分配被定义为鲁棒的。此外,必须在数学上量化资源分配的鲁棒性程度,例如,有多少台机器可能发生故障,在违反性能要求之前,系统参数的估计可能有多不准确? 具体而言,本研究解决的设计:数学上精确和广泛适用的技术建模和量化的鲁棒性的资源分配对系统组件和环境条件的多重扰动。资源分配算法,其持续地计划和开发用于响应潜在故障、资源降级和系统环境中的其它变化的策略。 这项工作代表了大学和工业/政府实验室之间的合作伙伴关系,这些实验室致力于为工业和国防应用开发高可用性计算系统。 其成果将通过介绍、出版物、跨学科讲习班和技术转让广泛传播。
英文摘要
Parallel and distributed computing systems, consisting of a heterogeneous set of machines, software, and networks, frequently operate in environments where their performance degrades due to circumstances that change unpredictably, such as sudden machine failures or inaccuracies in the estimation of system parameters. An important question then arises: what extent of departure from the assumed circumstances will cause the performance to degrade to the point where the system cannot meet the specified requirements i.e., how robust is the system? The focus of this work is the design of methodologies for generating robustness metrics and using them in resource management. A resource allocation is defined to be robust if degradation in system performance remains within specified limits when certain perturbations in specified system parameters occur. Furthermore, a resource allocations degree of robustness must be mathematically quantified e.g., how many machines can fail, how inaccurate can estimates in system parameters be before a performance requirement violation occurs? Specifically, this research addresses the design of: mathematically precise and widely applicable techniques for modeling and quantifying the robustness of a resource allocation against multiple perturbations in system components and environmental conditions. resource allocation algorithms that continually plan and develop strategies for responding to potential faults, resource degradation, and other changes in system environment. This work represents a partnership between university and industry/government laboratories that are committed to developing high availability computing systems for industry and defense applications. Its results will be widely disseminated through presentations, publications, interdisciplinary workshops, and technology transfer.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MRI Collaborative Consortium: Acquisition of a Shared Supercomputer by the Rocky Mountain Advanced Computing Consortium
-
批准号:1532235
-
项目类别:Standard Grant
-
资助金额:$70.0万
-
财政年份:2015
-
负责人:Howard Siegel
-
依托单位:
CSR:Medium:Collaborative Research: Stochastically Robust Resource Allocation for Computing
-
批准号:0905399
-
项目类别:Standard Grant
-
资助金额:$104.25万
-
财政年份:2009
-
负责人:Howard Siegel
-
依托单位:
MRI: Acquisition of the ISTeC High Performance Computing Infrastructure for Science and Engineering Research Projects
-
批准号:0923386
-
项目类别:Standard Grant
-
资助金额:$62.73万
-
财政年份:2009
-
负责人:Howard Siegel
-
依托单位:
NSF/Purdue Workshop on Grand Challenges in Computer Architecture for the Support of High Performance Computing; Purdue University; December 11-13, 1991
-
批准号:9200735
-
项目类别:Standard Grant
-
资助金额:$2.91万
-
财政年份:1991
-
负责人:Howard Siegel
-
依托单位:
Infrastructure for Parallel Processing Research
-
批准号:9015696
-
项目类别:Continuing Grant
-
资助金额:$125.19万
-
财政年份:1991
-
负责人:Howard Siegel
-
依托单位:
国内基金
海外基金
强流低能加速器束流损失机理的Parallel PIC/MCC算法与实现
-
批准号:11805229
-
项目类别:青年科学基金项目
-
资助金额:27.0万元
-
批准年份:2018
-
负责人:张青鵾
-
依托单位: