ExaHDF5: Delivering Efficient Parallel I/O on Exascale Computing Systems

ExaHDF5: Delivering Efficient Parallel I/O on Exascale Computing Systems
复制标题

ExaHDF5:在百亿亿次计算系统上提供高效的并行 I/O

DOI:
10.1007/s11390-020-9822-9
复制
发表时间:
2020
影响因子:
1.9
通讯作者:
R. Warren
R. Warren
中科院分区:
计算机科学3区
文献类型:
--
作者:
S. Byna;S. Breitenfeld;Bin Dong;Q. Koziol;Elena Pourmal;Dana Robinson;Jérome Soumagne;Houjun Tang;V. Vishwanath;R. Warren

文献摘要

被引文献

相似文献

Exascale的科学应用产生和分析了大量数据。这些应用程序的关键要求是在Exascale系统上有效访问和管理此数据的能力。并行I/O,关键技术可以在计算节点和存储之间进行移动数据,面临着Exascale系统设计中考虑的新应用程序,内存和存储架构的巨大挑战。随着存储层次结构的扩展为包括节点 - 局部持续内存,突发缓冲区等以及基于磁盘的存储,这些层之间的数据移动必须有效。未来的并行I/O库应能够处理许多Terabytes及以后的文件大小。在本文中,我们描述了我们以分层数据格式(HDF5)的新功能(HDF5)是科学应用程序最流行的I/O库。 HDF5是用于在现有HPC系统上执行并行I/O的领导力计算设施中最常用的库之一。我们描述的最新功能包括:虚拟对象层(VOL),数据电梯,异步I/O,功能完整的单作者和多阅读器(完整的SWMR)以及并行查询。在本文中,我们介绍了这些功能,它们的实现以及应用程序和其他库的性能和功能好处。
Scientific applications at exascale generate and analyze massive amounts of data. A critical requirement of these applications is the capability to access and manage this data efficiently on exascale systems. Parallel I/O, the key technology enables moving data between compute nodes and storage, faces monumental challenges from new applications, memory, and storage architectures considered in the designs of exascale systems. As the storage hierarchy is expanding to include node-local persistent memory, burst buffers, etc., as well as disk-based storage, data movement among these layers must be efficient. Parallel I/O libraries of the future should be capable of handling file sizes of many terabytes and beyond. In this paper, we describe new capabilities we have developed in Hierarchical Data Format version 5 (HDF5), the most popular parallel I/O library for scientific applications. HDF5 is one of the most used libraries at the leadership computing facilities for performing parallel I/O on existing HPC systems. The state-of-the-art features we describe include: Virtual Object Layer (VOL), Data Elevator, asynchronous I/O, full-featured single-writer and multiple-reader (Full SWMR), and parallel querying. In this paper, we introduce these features, their implementations, and the performance and feature benefits to applications and other libraries.