From Bare Metal to Virtual: Lessons Learned when a Supercomputing Institute Deploys its First Cloud

From Bare Metal to Virtual: Lessons Learned when a Supercomputing Institute Deploys its First Cloud
复制标题

从裸机到虚拟:超级计算机构部署首个云时的经验教训

DOI:
--
复制
发表时间:
2018
期刊:
Practice and Experience in Advanced Research Computing
影响因子:
--
通讯作者:
James C. Wilgenbusch
James C. Wilgenbusch
中科院分区:
--
文献类型:
--
作者:
Evan F. Bollig;James C. Wilgenbusch

文献摘要

被引文献

相似文献

作为明尼苏达大学研究计算服务的主要提供商,明尼苏达超级计算所(MSI)长期以来一直负责满足数以千计的用户群的需求。近年来,MSI--与许多其他HPC中心一样-观察到对自助式、按需、数据密集型研究的需求日益增长,以及许多用于研究目的的新的受控访问数据集的出现。有鉴于此,MSI构建了一个新的内部部署云服务,名为Stratus,它的架构从头开始设计,以轻松满足数据使用协议并填补传统HPC留下的四个空白。由此产生的OpenStack云由HPC特定的计算节点构建,并由Cave存储支持,旨在完全遵守NIH基因组数据共享政策规定的控制。在这里,我们介绍了在雄心勃勃的冲刺过程中学到的12条经验教训,在不到18个月的时间里,Stratus从开始到投入生产。这一时间表的重要组成部分包括发展新的领导角色、工作人员和用户培训以及用户支持文档,这些都很重要,但往往被忽视。在此过程中,所学到的经验教训远远超出了通常与获取、配置和维护大型系统相关的技术挑战。
As primary provider for research computing services at the University of Minnesota, the Minnesota Supercomputing Institute (MSI) has long been responsible for serving the needs of a user-base numbering in the thousands. In recent years, MSI---like many other HPC centers---has observed a growing need for self-service, on-demand, data-intensive research, as well as the emergence of many new controlled-access datasets for research purposes. In light of this, MSI constructed a new on-premise cloud service, named Stratus, which is architected from the ground up to easily satisfy data-use agreements and fill four gaps left by traditional HPC. The resulting OpenStack cloud, constructed from HPC-specific compute nodes and backed by Ceph storage, is designed to fully comply with controls set forth by the NIH Genomic Data Sharing Policy. Herein, we present twelve lessons learned during the ambitious sprint to take Stratus from inception and into production in less than 18 months. Important, and often overlooked, components of this timeline included the development of new leadership roles, staff and user training, and user support documentation. Along the way, the lessons learned extended well beyond the technical challenges often associated with acquiring, configuring, and maintaining large-scale systems.