Running a Pre-exascale, Geographically Distributed, Multi-cloud Scientific Simulation

Running a Pre-exascale, Geographically Distributed, Multi-cloud Scientific Simulation
复制标题

DOI:
10.1007/978-3-030-50743-5_2
复制
发表时间:
2020-05-22
期刊:
High Performance Computing
影响因子:
--
通讯作者:
Schultz D
Schultz D
中科院分区:
其他
文献类型:
--
作者:
Sfiligoi I;Würthwein F;Riedel B;Schultz D

文献摘要

被引文献

相似文献

当我们接近Exascale时代时,重要的是要验证现有的框架和工具是否仍然可以在这种规模下工作。此外,公共云计算已经成为原型设计和紧急计算的可行解决方案。利用云的弹性,我们已经建立了一个pre-exascale HTCondor设置,用于在云中运行科学模拟,选择的应用程序是IceCube的光子传播模拟。也就是说,这不是一个纯粹的示范运行,但它也被用来为冰立方合作产生有价值的和急需的科学结果。为了达到预期的规模,我们在亚马逊网络服务、微软Azure和谷歌云平台的许多地理区域的8种GPU模型上聚合了GPU资源。使用这种设置,我们达到了超过51k GPU的峰值,对应于近380个pflop32,总的集成计算约为100k GPU小时。在本文中,我们提供了设置的描述,发现和克服的问题,以及练习的实际科学输出的简短描述。
As we approach the Exascale era, it is important to verify that the existing frameworks and tools will still work at that scale. Moreover, public Cloud computing has been emerging as a viable solution for both prototyping and urgent computing. Using the elasticity of the Cloud, we have thus put in place a pre-exascale HTCondor setup for running a scientific simulation in the Cloud, with the chosen application being IceCube’s photon propagation simulation. I.e. this was not a purely demonstration run, but it was also used to produce valuable and much needed scientific results for the IceCube collaboration. In order to reach the desired scale, we aggregated GPU resources across 8 GPU models from many geographic regions across Amazon Web Services, Microsoft Azure, and the Google Cloud Platform. Using this setup, we reached a peak of over 51k GPUs corresponding to almost 380 PFLOP32s, for a total integrated compute of about 100k GPU hours. In this paper we provide the description of the setup, the problems that were discovered and overcome, as well as a short description of the actual science output of the exercise.