Management of an academic HPC cluster: The UL experience

Management of an academic HPC cluster: The UL experience
复制标题

学术 HPC 集群的管理:UL 经验

DOI:
--
复制
发表时间:
2014
期刊:
International Symposium on High Performance Computing Systems and Applications
影响因子:
--
通讯作者:
F. Georgatos
F. Georgatos
中科院分区:
--
文献类型:
--
作者:
Sébastien Varrette;P. Bouvry;Hyacinthe Cartiaux;F. Georgatos

文献摘要

被引文献

相似文献

处理能力、数据存储和传输能力的密集增长已经彻底改变了科学的许多方面。这些资源对于在许多应用领域取得高质量结果至关重要。在此背景下,卢森堡大学 (UL) 自 2007 年起由一个非常小的团队运营高性能计算 (HPC) 设施和相关存储。桥接计算和存储方面是 UL 服务的要求 - 原因既合法(某些数据可能不会移动)又与性能相关。如今,来自 UL 三个院系和/或两个跨学科中心的人员都是该设施的用户。更具体地说,系统生物医学(由 LCSB 负责)和安全、可靠性和信任(由 SnT 负责)等关键研究重点需要访问此类 HPC 设施,以便在适当的环境中发挥作用。 HPC 解决方案的管理是一项复杂的工作,也是一个需要不断讨论和改进的领域。 UL HPC 设施和衍生的部署服务是一个复杂的计算系统,按其规模进行管理:在撰写本文时,它由 150 台服务器、368 个节点(3880 个计算核心)和 1996 TB 共享存储组成,仅由三个人使用基于 Puppet [1]、FAI [2] 和 Capistrano [3] 的先进 IT 自动化解决方案进行配置、监控和操作。本文涵盖了与管理如此复杂的基础设施相关的所有方面,无论是技术方面还是行政方面。大多数设计选择或实施的方法都是由多年解决研究需求的经验推动的,主要是在 HPC 领域,但也包括补充服务(通常基于 Web)。在这样的背景下,我们试图以灵活便捷的方式解答很多技术问题。这份经验报告可能会引起属于公共或私营部门的其他研究中心和大学的兴趣,这些研究中心和大学在集群架构和管理方面寻找良好的(即使不是最佳的)实践。
The intensive growth of processing power, data storage and transmission capabilities has revolutionized many aspects of science. These resources are essential to achieve high-quality results in many application areas. In this context, the University of Luxembourg (UL) operates since 2007 an High Performance Computing (HPC) facility and the related storage by a very small team. The aspect of bridging computing and storage is a requirement of UL service - the reasons are both legal (certain data may not move) and performance related. Nowadays, people from the three faculties and/or the two Interdisciplinary centers within the UL, are users of this facility. More specifically, key research priorities such as Systems Bio-medicine (by LCSB) and Security, Reliability & Trust (by SnT) require access to such HPC facilities in order to function in an adequate environment. The management of HPC solutions is a complex enterprise and a constant area for discussion and improvement. The UL HPC facility and the derived deployed services is a complex computing system to manage by its scale: at the moment of writing, it consists of 150 servers, 368 nodes (3880 computing cores) and 1996 TB of shared storage which are all configured, monitored and operated by only three persons using advanced IT automation solutions based on Puppet [1], FAI [2] and Capistrano [3]. This paper covers all the aspects in relation to the management of such a complex infrastructure, whether technical or administrative. Most design choices or implemented approaches have been motivated by several years of experience in addressing research needs, mainly in the HPC area but also in complementary services (typically Web-based). In this context, we tried to answer in a flexible and convenient way many technological issues. This experience report may be of interest for other research centers and universities belonging either to the public or the private sector looking for good if not best practices in cluster architecture and management.