Running HPC applications on many million cores Cloud

Running HPC applications on many million cores Cloud
复制标题

在数百万个核心上运行 HPC 应用程序 云

DOI:
--
复制
发表时间:
2017
期刊:
International Convention on Information and Communication Technology, Electronics and Microelectronics
影响因子:
--
通讯作者:
D. Ogrizovic
D. Ogrizovic
中科院分区:
--
文献类型:
--
作者:
D. Tomić;Z. Car;D. Ogrizovic

文献摘要

被引文献

相似文献

尽管云架构在硬件和软件方面进行了多方面的改进,但在运行高性能通信密集型应用时,商用超级计算机与云仍然存在巨大的性能差距。为了找出阻碍他们在云上更好扩展的原因,我们在HPE OpenStack试验床上评估了HPL和NAMD基准,并在Rijeka大学超级计算中心的超级计算机上评估了NAMD基准。我们的结果揭示了两个主要瓶颈:互联的吞吐量和云编排层,以及负责管理云实例之间的通信的其他层。我们调查了抖动的影响,但没有发现对性能的显著影响。我们的结论是,仅通过增加互连吞吐量并不能提高云中HPC通信密集型HPC应用的可扩展性。这也得到了惠普实验室执行的NAMD和圣地亚哥超级计算中心执行的HPL基准测试的支持。我们提出了两种可能的可伸缩性改进方案。一种采用分布式云协调层模型;另一种采用裸机容器。如果我们希望看到HPC应用扩展到超过数百万个云核心,那么高效的负载平衡仍然是必须的。为此,我们提出了一种新的基于SLEM的负载均衡策略。
Despite the various hardware and software improvements in Cloud architecture, there still exists the huge performance gap between the commodity supercomputers and Cloud when running HPC communication intensive applications. In order to find what is preventing them to better scale on Cloud, we evaluated HPL and NAMD benchmarks on HPE Openstack testbed, and NAMD benchmarks on supercomputer located at Rijeka University Supercomputing Center. Our results revealed two major bottlenecks: the throughput of the interconnect, and Cloud orchestration layer, among other responsible for the management of the communication between Cloud instances. We investigated the influence of jittering, but did not find the significant influence on performance. Our conclusion is that by solely increasing the interconnect throughput, one will not improve the scalability of HPC communication intensive HPC applications in Cloud. This is also backed up with NAMD performed at HP Labs, and with HPL benchmark performed at San Diego Supercomputing Center. We propose two possible scenarios of scalability improvements. One with distributed model of Cloud Orchestration layer; another with bare metal containers. Efficient load balancing remains the must if we want to see HPC applications scaling over many million Cloud cores. For this, we propose novel SLEM based load balancing strategy.