SPRUCE: A System for Supporting Urgent High-Performance Computing

SPRUCE: A System for Supporting Urgent High-Performance Computing
复制标题

SPRUCE:支持紧急高性能计算的系统

DOI:
--
复制
发表时间:
2006
期刊:
Grid-Based Problem Solving Environments
影响因子:
--
通讯作者:
Ivan Beschastnikh
Ivan Beschastnikh
中科院分区:
--
文献类型:
--
作者:
P. Beckman;S. Nadella;N. Trebon;Ivan Beschastnikh

文献摘要

被引文献

相似文献

使用高性能计算的建模和仿真在决策和预测中发挥着越来越重要的作用。对于时间关键的紧急决策支持应用程序,如流感建模和恶劣天气预测,晚的结果可能是无用的。需要专门的基础设施来快速提供计算资源。本文介绍了SPRUCE的体系结构和实现,系统支持紧急计算在传统的超级计算机和分布式计算网格。目前部署在TeraGrid上的SPRUCE为用户提供了“通行权令牌”,在紧急计算需要的情况下,可以从基于Web的门户或Web服务调用激活这些令牌。令牌是可转移的,并且可以限制到特定的资源集和优先级。一旦会话被激活,作业提交可能会请求提升的优先级。基于本地策略,计算资源可以例如通过抢占活动作业或提高作业在队列中的优先级来进行响应。本文还探讨了SPRUCE架构和基于令牌的激活紧急计算应用程序的优点和缺点。
Modeling and simulation using high-performance computing are playing an increasingly important role in decision making and prediction. For time-critical emergency decision support applications, such as influenza modeling and severe weather prediction, late results may be useless. A specialized infrastructure is needed to provide computational resources quickly. This paper describes the architecture and implementation of SPRUCE, a system for supporting urgent computing on both traditional supercomputers and distributed computing Grids. Currently deployed on the TeraGrid, SPRUCE provides users with “right-of-way tokens” that can be activated from a Web-based portal or Web service invocation in the event of an urgent computing need. Tokens are transferrable and can be restricted to specific resource sets and priority levels. Once a session is activated, job submissions may request elevated priority. Based on local policy, computing resources can respond, for example, by preempting active jobs or raising the job’s priority in the queue. This paper also explores the strengths and weaknesses of the SPRUCE architecture and token-based activation for urgent computing applications.