Trestles: a high-productivity HPC system targeted to modest-scale and gateway users

Trestles: a high-productivity HPC system targeted to modest-scale and gateway users
复制标题

Trestles:面向中等规模和网关用户的高生产率 HPC 系统

DOI:
--
复制
发表时间:
2011
期刊:
TeraGrid Conference
影响因子:
--
通讯作者:
William S. Young
William S. Young
中科院分区:
--
文献类型:
--
作者:
Richard L. Moore;David L. Hart;W. Pfeiffer;M. Tatineni;Kenneth Yoshimoto;William S. Young

文献摘要

被引文献

相似文献

Trestles是SDSC的一种新的100TF HPC资源,旨在提高TeraGRID中适度尺度和网关用户的科学生产力。本文讨论了对该用户群的目标以及计划的操作策略和程序以优化科学生产力,包括针对科学生产力的基本原理,包括除了传统的系统利用外,还重新集中在周转时间上。 teragrid用户的一部分出奇的一小部分运行适度的作业(例如<1k核心),而越来越多的TeraGrid用户可以通过网关访问HPC资源;尽管这些用户代表了用户群的很大一部分,但他们消耗了较小的Teragrid资源。因此,尽管Trestles并不是Teragrid中最大的HPC资源,但它将能够在旨在提高其生产率的环境中支持这一大型Teragrid用户。该针对性的使用模型还可以释放其他需要大规模,SMP或其他特定资源功能的用户/作业的Teragrid系统。栈桥的主要区别者之一是,它将被分配和计划以优化队列等待时间和扩展因子,以及传统的系统利用度量。此外,带有32个内核和64GB DRAM的节点设计将容纳许多没有节点通信的工作,而120GB本地闪存将加快许多应用程序。系统上安装了一套强大的应用程序软件,包括高斯,Blast,Abaqus,Gamess,Amber和NAMD。标准工作限制为32个节点(1K内核)和48小时的运行时间,但可以做出例外,尤其是长达2周的长期工作。常规系统预订确保始终为较短,较小的工作和可提供用户的预订提供一些节点,以确保用户可以预测对系统的访问。可以在独家或共享模式下访问节点。最后,Trestles是唯一具有自动按需访问的Teragrid资源。有限数量的节点被配置为工作以“处于风险的危险”(收取的使用率折扣),并受到按需工作的预先杀害(这些工作率有溢价)。随着使用模式的出现和用户提供反馈以进一步提高其生产率,分配,调度和软件环境将随着时间的流逝而随着时间的推移进行调整和调整。
Trestles is a new 100TF HPC resource at SDSC designed to enhance scientific productivity for modest-scale and gateway users within the TeraGrid. This paper discusses the Trestles hardware and user environment, as well as the rationale for targeting this user base and the planned operational policies and procedures to optimize scientific productivity, including a focus on turnaround time in addition to the traditional system utilization. A surprisingly large fraction of TeraGrid users run modest-scale jobs (e.g. <1K cores), and an increasing fraction of TeraGrid users access HPC resources via gateways; while these users represent a large percentage of the user base, they consume a smaller fraction of the TeraGrid resources. Thus, while Trestles is not the largest HPC resource in TeraGrid, it will be able to support this large class of TeraGrid users in an environment designed to enhance their productivity. This targeted usage model also frees up other TeraGrid systems for users/jobs that require large-scale, SMP or other specific resource features. One of the key differentiators for Trestles is that it will be allocated and scheduled to optimize queue wait times and expansion factors, as well as the traditional system utilization metric. In addition, the node design, with 32 cores and 64GB DRAM, will accommodate many jobs without inter-node communications, while the 120GB local flash memory will speed up many applications. A robust set of application software, including Gaussian, BLAST, Abaqus, GAMESS, Amber and NAMD, is installed on the system. Standard job limits are 32 nodes (1K cores) and 48 hours runtime, but exceptions can be made, particularly for long jobs up to 2 weeks. Standing system reservations ensure that some nodes are always set aside for shorter, smaller jobs, and user-settable reservations are available to ensure users predictable access to the system. Nodes can be accessed in exclusive or shared mode. Finally, Trestles is the only TeraGrid resource with automatic on-demand access; a limited number of nodes is configured for jobs to "run at risk" (with a discount in the usage rate charged) and be subject to being pre-emptively killed by on-demand jobs (which carry a premium in the usage rate). The allocation, scheduling and software environments will be adjusted and tuned over time as usage patterns emerge and users provide feedback to further enhance their productivity.