Developing Distributed High-performance Computing Capabilities of an Open Science Platform for Robust Epidemic Analysis

Developing Distributed High-performance Computing Capabilities of an Open Science Platform for Robust Epidemic Analysis
复制标题

开发开放科学平台的分布式高性能计算能力以进行稳健的流行病分析

DOI:
10.1109/ipdpsw59300.2023.00143
复制
发表时间:
2023
期刊:
2023 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW
影响因子:
--
通讯作者:
Ozik, Jonathan
Ozik, Jonathan
中科院分区:
--
文献类型:
--
作者:
Collier, Nicholson;Wozniak, Justin M.;Stevens, Abby;Babuji, Yadu;Binois, Mickaël;Fadikar, Arindam;Würth, Alexandra;Chard, Kyle;Ozik, Jonathan

文献摘要

参考文献

被引文献

相似文献

COVID-19对科学合作产生了前所未有的影响。这场大流行及其在科学界的广泛响应,在领域专家、数学建模者和科学计算专家之间建立了新的关系。然而,在计算方面,它也揭示了研究人员利用先进计算系统的能力方面的关键差距。这些具有挑战性的领域包括访问可扩展的计算系统,将模型和工作流程移植到新系统,共享不同大小的数据,以及产生可以由他人复制和验证的结果。根据我们团队在COVID-19大流行期间支持公共卫生决策者的工作,以及在将高性能计算(HPC)应用于复杂社会系统建模方面所发现的能力差距,我们提出了OSPREY的目标,要求和初步实施,OSPREY是一个强大的流行病分析开放科学平台。原型实施展示了一个集成的、算法驱动的HPC工作流架构,协调联合HPC资源之间的任务,并对每个资源进行可靠、安全和自动化的访问。我们展示了可扩展和容错的任务执行,异步API,以支持快速的时间解决方案的算法,一个包容性,多语言的方法,和有效的广域数据管理。示例OSPREY代码在公共存储库中可用。
COVID-19 had an unprecedented impact on scientific collaboration. The pandemic and its broad response from the scientific community has forged new relationships among domain experts, mathematical modelers, and scientific computing specialists. Computationally, however, it also revealed critical gaps in the ability of researchers to exploit advanced computing systems. These challenging areas include gaining access to scalable computing systems, porting models and workflows to new systems, sharing data of varying sizes, and producing results that can be reproduced and validated by others. Informed by our team's work in supporting public health decision makers during the COVID-19 pandemic and by the identified capability gaps in applying high-performance computing (HPC) to the modeling of complex social systems, we present the goals, requirements, and initial implementation of OSPREY, an open science platform for robust epidemic analysis. The prototype implementation demonstrates an integrated, algorithm-driven HPC workflow architecture, coordinating tasks across federated HPC resources, with robust, secure and automated access to each of the resources. We demonstrate scalable and fault-tolerant task execution, an asynchronous API to support fast time-to-solution algorithms, an inclusive, multi-language approach, and efficient wide-area data management. The example OSPREY code is made available on a public repository.
DOI: 10.1007/s11192-021-03873-7
发表时间: 2021
期刊: Scientometrics
影响因子: 3.9
作者:
Cai X;Fry CV;Wagner CS
通讯作者: Wagner CS
ResearchOps:科学应用中的 DevOps 案例
DOI: --
发表时间: 2015
期刊: IFIP/IEEE Symposium on Integrated Network Management
影响因子: --
作者:
M. D. Bayser;L. Azevedo;Renato F. G. Cerqueira
通讯作者: Renato F. G. Cerqueira
DOI: 10.4016/3366.01
发表时间: 2007
期刊: --
影响因子: --
作者:
M. Feller;Ian T Foster;Stuart Martin
通讯作者: M. Feller;Ian T Foster;Stuart Martin
DOI: 10.1109/mlhpc54614.2021.00007
发表时间: 2021-10
期刊: 2021 IEEE/ACM Workshop on Machine Learning in High Performance Computing Environments (MLHPC)
影响因子: --
作者:
Logan T. Ward;G. Sivaraman;J. G. Pauloski;Y. Babuji;Ryan Chard;Naveen K. Dandu;P. Redfern;R. Assary;K. Chard;L. Curtiss;R. Thakur;Ian T. Foster
通讯作者: Logan T. Ward;G. Sivaraman;J. G. Pauloski;Y. Babuji;Ryan Chard;Naveen K. Dandu;P. Redfern;R. Assary;K. Chard;L. Curtiss;R. Thakur;Ian T. Foster
DOI: 10.1109/mcse.2019.2919690
发表时间: 2019-07-01
影响因子: 2.1
作者:
Deelman, Ewa;Vahi, Karan;Livny, Miron
通讯作者: Livny, Miron