Robust Parallel and Distributed Computing Systems
Robust Parallel and Distributed Computing Systems
批准号:
0615170
负责人:
Howard Siegel
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-06-15 至 2010-05-31
中文摘要
并行和分布式计算系统由一组异构的机器、软件和网络组成,它们经常在由于不可预测的变化而导致性能下降的环境中运行,例如突然的机器故障或对系统参数的估计不准确。一个重要的问题随之而来:偏离假设环境的程度会导致性能下降到系统无法满足指定需求的程度,也就是说,系统有多健壮?这项工作的重点是设计用于生成健壮性度量并在资源管理中使用它们的方法。如果在特定系统参数发生某些扰动时,系统性能的退化保持在规定的范围内,则资源分配被定义为鲁棒的。此外,鲁棒性的资源分配程度必须在数学上量化,例如,有多少机器会失败,在发生性能需求冲突之前对系统参数的估计有多不准确?具体来说,本研究解决了:数学上精确和广泛适用的技术的设计,用于建模和量化资源分配对系统组件和环境条件的多重扰动的鲁棒性。不断规划和开发策略以响应系统环境中的潜在故障、资源退化和其他变化的资源分配算法。这项工作代表了大学和工业/政府实验室之间的合作关系,致力于为工业和国防应用开发高可用性计算系统。其成果将通过演讲、出版物、跨学科讲习班和技术转让等方式广泛传播。
英文摘要
Parallel and distributed computing systems, consisting of a heterogeneous set of machines, software, and networks, frequently operate in environments where their performance degrades due to circumstances that change unpredictably, such as sudden machine failures or inaccuracies in the estimation of system parameters. An important question then arises: what extent of departure from the assumed circumstances will cause the performance to degrade to the point where the system cannot meet the specified requirements i.e., how robust is the system? The focus of this work is the design of methodologies for generating robustness metrics and using them in resource management. A resource allocation is defined to be robust if degradation in system performance remains within specified limits when certain perturbations in specified system parameters occur. Furthermore, a resource allocations degree of robustness must be mathematically quantified e.g., how many machines can fail, how inaccurate can estimates in system parameters be before a performance requirement violation occurs? Specifically, this research addresses the design of: mathematically precise and widely applicable techniques for modeling and quantifying the robustness of a resource allocation against multiple perturbations in system components and environmental conditions. resource allocation algorithms that continually plan and develop strategies for responding to potential faults, resource degradation, and other changes in system environment. This work represents a partnership between university and industry/government laboratories that are committed to developing high availability computing systems for industry and defense applications. Its results will be widely disseminated through presentations, publications, interdisciplinary workshops, and technology transfer.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MRI Collaborative Consortium: Acquisition of a Shared Supercomputer by the Rocky Mountain Advanced Computing Consortium
-
批准号:1532235
-
项目类别:Standard Grant
-
资助金额:$70.0万
-
财政年份:2015
-
负责人:Howard Siegel
-
依托单位:
CSR:Medium:Collaborative Research: Stochastically Robust Resource Allocation for Computing
-
批准号:0905399
-
项目类别:Standard Grant
-
资助金额:$104.25万
-
财政年份:2009
-
负责人:Howard Siegel
-
依托单位:
MRI: Acquisition of the ISTeC High Performance Computing Infrastructure for Science and Engineering Research Projects
-
批准号:0923386
-
项目类别:Standard Grant
-
资助金额:$62.73万
-
财政年份:2009
-
负责人:Howard Siegel
-
依托单位:
NSF/Purdue Workshop on Grand Challenges in Computer Architecture for the Support of High Performance Computing; Purdue University; December 11-13, 1991
-
批准号:9200735
-
项目类别:Standard Grant
-
资助金额:$2.91万
-
财政年份:1991
-
负责人:Howard Siegel
-
依托单位:
Infrastructure for Parallel Processing Research
-
批准号:9015696
-
项目类别:Continuing Grant
-
资助金额:$125.19万
-
财政年份:1991
-
负责人:Howard Siegel
-
依托单位:
国内基金
海外基金
强流低能加速器束流损失机理的Parallel PIC/MCC算法与实现
-
批准号:11805229
-
项目类别:青年科学基金项目
-
资助金额:27.0万元
-
批准年份:2018
-
负责人:张青鵾
-
依托单位: