Stimulus: Accelerate Data Management for Scientific AI applications in HPC

Stimulus: Accelerate Data Management for Scientific AI applications in HPC
复制标题

DOI:
10.1109/ccgrid54584.2022.00020
复制
发表时间:
2022-05
期刊:
2022 22nd IEEE International Symposium on Cluster, Cloud and Internet Computing (CCGrid)
影响因子:
--
通讯作者:
H. Devarajan;Anthony Kougkas;Huihuo Zheng;V. Vishwanath;Xian-He Sun
H. Devarajan;Anthony Kougkas;Huihuo Zheng;V. Vishwanath;Xian-He Sun
中科院分区:
其他
文献类型:
--
作者:
H. Devarajan;Anthony Kougkas;Huihuo Zheng;V. Vishwanath;Xian-He Sun

文献摘要

被引文献

相似文献

现代科学工作流程将模拟与人工智能分析相结合,通过频繁交换数据来加快科学研究的时间,从而降低模拟平面的复杂性。然而,由于人工智能框架中缺乏对科学数据格式的支持,这种数据交换在性能和可移植性方面受到限制。我们需要一个内聚机制来有效地将大规模复杂的科学数据格式(如HDF5、PnetCDF、ADIOS2、GNCF和Silo)集成到流行的AI框架(如TensorFlow、PyTorch和Caffe)中。为此,我们设计了一个数据管理库Stimulus,用于将科学数据有效地摄取到流行的AI框架中。我们利用StimOps函数和StimPack抽象来实现科学数据格式与任何AI框架的集成。评估表明,Stimulus在不同用例下的几个大型应用程序,如Cosmic Tagger(在PyTorch中消费HDF5数据集),Distributed FFN(在TensorFlow中消费HDF5数据集)和CosmoFlow(将HDF5转换为TFRecord,然后在TensorFlow中消费)的性能分别优于5.3倍,2.9倍和1.9倍,理想的I/O可扩展性在Summit超级计算机上高达768个gpu。通过Stimulus,我们可以移植地扩展现有流行的AI框架,以凝聚支持任何复杂的科学数据格式,并有效地扩展大型超级计算机上的应用程序。
Modern scientific workflows couple simulations with AI-powered analytics by frequently exchanging data to accelerate time-to-science to reduce the complexity of the simulation planes. However, this data exchange is limited in performance and portability due to a lack of support for scientific data formats in AI frameworks. We need a cohesive mechanism to effectively integrate at scale complex scientific data formats such as HDF5, PnetCDF, ADIOS2, GNCF, and Silo into popular AI frameworks such as TensorFlow, PyTorch, and Caffe. To this end, we designed Stimulus, a data management library for ingesting scientific data effectively into the popular AI frameworks. We utilize the StimOps functions along with StimPack abstraction to enable the integration of scientific data formats with any AI framework. The evaluations show that Stimulus outperforms several large-scale applications with different use-cases such as Cosmic Tagger (consuming HDF5 dataset in PyTorch), Distributed FFN (consuming HDF5 dataset in TensorFlow), and CosmoFlow (converting HDF5 into TFRecord and then consuming that in TensorFlow) by 5.3 x, 2.9 x, and 1.9 x respectively with ideal I/O scalability up to 768 GPUs on the Summit supercomputer. Through Stimulus, we can portably extend existing popular AI frameworks to cohesively support any complex scientific data format and efficiently scale the applications on large-scale supercomputers.