Demonstrating a Pre-Exascale, Cost-Effective Multi-Cloud Environment for Scientific Computing: Producing a fp32 ExaFLOP hour worth of IceCube simulation data in a single workday

Demonstrating a Pre-Exascale, Cost-Effective Multi-Cloud Environment for Scientific Computing: Producing a fp32 ExaFLOP hour worth of IceCube simulation data in a single workday
复制标题

演示用于科学计算的前百亿亿次规模、经济高效的多云环境:在单个工作日内生成 fp32 ExaFLOP 小时的 IceCube 模拟数据

DOI:
10.1145/3311790.3396625
复制
发表时间:
2020
期刊:
PEARC '20: Practice and Experience in Advanced Research Computing
影响因子:
--
通讯作者:
Brik, Vladimir
Brik, Vladimir
中科院分区:
--
文献类型:
--
作者:
Sfiligoi, Igor;Schultz, David;Riedel, Benedikt;Wuerthwein, Frank;Barnet, Steve;Brik, Vladimir

文献摘要

参考文献

被引文献

相似文献

科学计算需求随着时间的推移而急剧增长,并且在以前不是计算密集型的科学领域中不断扩展。当计算工作流的峰值远远超过其本地计算资源的容量时,应该从其他地方临时提供容量,以满足最后期限并增加科学产出。公共云已经成为一个有吸引力的选择,因为它们能够在最小的提前通知下进行配置。成本效益高的实例的可用容量尚未得到充分了解。本文介绍了使用从三大云提供商(即Amazon Web Services,Microsoft Azure和Google Cloud Platform)收集的具有成本效益的GPU实例以抢占模式扩展IceCube的生产HTCondor池。使用这种设置,我们在整个工作日维持了大约15k GPU,相当于大约170个PFLOP32,集成了超过一个EFLOP32小时的科学输出,价格约为6万美元。在本文中,我们提供了云实例选择背后的推理,对设置的描述和对所提供资源的分析,以及对练习的实际科学输出的简短描述。
Scientific computing needs are growing dramatically with time and are expanding in science domains that were previously not compute intensive. When compute workflows spike well in excess of the capacity of their local compute resource, capacity should be temporarily provisioned from somewhere else to both meet deadlines and to increase scientific output. Public Clouds have become an attractive option due to their ability to be provisioned with minimal advance notice. The available capacity of cost-effective instances is not well understood. This paper presents expanding the IceCube's production HTCondor pool using cost-effective GPU instances in preemptible mode gathered from the three major Cloud providers, namely Amazon Web Services, Microsoft Azure and the Google Cloud Platform. Using this setup, we sustained for a whole workday about 15k GPUs, corresponding to around 170 PFLOP32s, integrating over one EFLOP32 hour worth of science output for a price tag of about $60k. In this paper, we provide the reasoning behind Cloud instance selection, a description of the setup and an analysis of the provisioned resources, as well as a short description of the actual science output of the exercise.
利用 Google 云平台进行按需紧急高性能计算
DOI: 10.1109/urgenthpc49580.2019.00008
发表时间: 2019
期刊: 2019 IEEE/ACM HPC for Urgent Decision Making (UrgentHPC
影响因子: --
作者:
Posey, Brandon;Deer, Adam;Gorman, Wyatt;July, Vanessa;Kanhere, Neeraj;Speck, Dan;Wilson, Boyd;Apon, Amy
通讯作者: Apon, Amy
OASIS:开放科学网格的数据和软件分发服务
DOI: 10.1088/1742-6596/513/3/032013
发表时间: 2014
期刊: Journal of Physics: Conference Series
影响因子: --
作者:
B. Bockelman;J. C. Bejar;J. D. Stefano;J. Hover;Robert Quick;S. Teige
通讯作者: S. Teige
使用 GlideinWMS 进行云爆发:满足科学工作流程不断增长的计算需求的方法
DOI: --
发表时间: 2014
期刊:
影响因子: --
作者:
P. Mhashilkar;A. Tiradani;B. Holzman;K. Larson;I. Sfiligoi;M. Rynge
通讯作者: M. Rynge
DOI: 10.1007/978-3-030-50743-5_2
发表时间: 2020-05-22
期刊: High Performance Computing
影响因子: --
作者:
Sfiligoi I;Würthwein F;Riedel B;Schultz D
通讯作者: Schultz D