Efficient and Automated Deployment Architecture for OpenStack in TianHe SuperComputing Environment

Efficient and Automated Deployment Architecture for OpenStack in TianHe SuperComputing Environment
复制标题

DOI:
10.1109/tpds.2021.3127128
复制
发表时间:
2022-08
影响因子:
5.3
通讯作者:
Bingting Jiang;Zhuo Tang;X. Xiao;Jing Yao;Ronghui Cao;Kenli Li
Bingting Jiang;Zhuo Tang;X. Xiao;Jing Yao;Ronghui Cao;Kenli Li
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bingting Jiang;Zhuo Tang;X. Xiao;Jing Yao;Ronghui Cao;Kenli Li

文献摘要

相似文献

近年来,随着全球金融危机和公共安全事件(如COVID-19)的大规模爆发,高性能计算被广泛应用于风险预测、疫苗研发等领域。在高性能计算基础设施应对计算需求瞬间爆炸的场景下,一个关键问题是通过快速构建计算集群,提供大规模的计算能力灵活分配和调整。现有的大规模计算集群部署解决方案通常利用源代码部署或其他部署工具。现有部署方法的最大挑战是减少过多的映像分发时间并避免配置缺陷。本文设计了一种基于OpenStack云平台的智能分布式注册中心部署(IDRD)架构,通过多个注册中心的容器化部署,自适应地放置分布式镜像仓库。提出了一种服务器负载优先算法来解决IDRD中的多注册表放置问题。在此基础上,设计了一种基于需求密度的分簇算法,该算法能够优化IDRD的全局性能,提高大规模集群的负载均衡能力,并在天河超级计算环境中得到了实现。大量实验结果表明,IDRD可以有效地减少组件映像分发时间的30%~50%,显著提高大规模集群部署的效率.
Recently, with the large-scale outbreak of the global financial crisis and public safety incidents (such as COVID-19), high-performance computing has been widely applied to risk prediction, vaccine development, and other fields. In scenarios where high-performance computing infrastructure responds to the instantaneous explosion of computing demands, a crucial issue is to provide large-scale flexible allocation and adjustment of computing capability by rapidly constructing computing clusters. Existing large-scale computing cluster deployment solutions usually utilize source code deployment or other deployment tools. The great challenge of existing deployment methods is to reduce excessive image distribution time and refrain from configuration defects. In this article, we design an intelligent distributed registry deployment (IDRD) architecture based on the OpenStack cloud platform, which adaptively places distributed image repositories using the containerized deployment of multiple registries. We propose a server load priority algorithm to solve multiple registries placement problems in IDRD. Furthermore, we devise a clustering algorithm based on demand density that can optimize the global performance of IDRD and improve large-scale cluster load balancing capabilities, which has been implemented in the TianHe Supercomputing environment. Extensive experimental results demonstrate that IDRD can effectively reduce $30\%$30%-$50\%$50% of the distribution time of component images and significantly improve the efficiency of large-scale cluster deployment.