Toward Transparent Data Management in Multi-Layer Storage Hierarchy of HPC Systems

Toward Transparent Data Management in Multi-Layer Storage Hierarchy of HPC Systems
复制标题

DOI:
10.1109/ic2e.2018.00046
复制
发表时间:
2018-05
期刊:
2018 IEEE International Conference on Cloud Engineering (IC2E)
影响因子:
--
通讯作者:
Bharti Wadhwa;S. Byna;A. Butt
Bharti Wadhwa;S. Byna;A. Butt
中科院分区:
其他
文献类型:
--
作者:
Bharti Wadhwa;S. Byna;A. Butt

文献摘要

被引文献

相似文献

即将推出的百亿亿次高性能计算 (HPC) 系统预计将包含多层存储层次结构,因此需要创新的存储和 I/O 机制。由于缺乏层次结构支持和语义接口,传统的基于磁盘和块的接口和文件系统在利用存储层次结构的能力方面面临着严峻的挑战。用于大规模系统上的科学数据管理的基于对象和语义丰富的数据抽象为这些挑战提供了可持续的解决方案。这种数据抽象还可以简化用户参与数据移动的过程。在本文中,我们迈出了实现此类对象抽象的第一步,并探索这些对象的存储机制以增强 I/O 性能,特别是对于科学应用程序。我们通过呈现来自两个现实世界 HPC 科学用例的数据 I/O 映射:等离子体物理模拟代码 (VPIC) 和宇宙学模拟代码 (HACC),探索基于对象的界面如何促进下一代可扩展计算系统。我们的存储模型将数据对象存储在不同的物理组织中,以支持跨内存/存储层次结构的数据移动。我们的实现可以很好地扩展到 16K 并行进程,并且与 MPI-IO 和 HDF5 等最先进的技术相比,我们基于对象的数据抽象和多级存储层次结构中的数据放置策略为科学数据实现了高达 7 倍的 I/O 性能改进。
Upcoming exascale high performance computing (HPC) systems are expected to comprise multi-tier storage hierarchy, and thus will necessitate innovative storage and I/O mechanisms. Traditional disk and block-based interfaces and file systems face severe challenges in utilizing capabilities of storage hierarchies due to the lack of hierarchy support and semantic interfaces. Object-based and semantically-rich data abstractions for scientific data management on large scale systems offer a sustainable solution to these challenges. Such data abstractions can also simplify users involvement in data movement. In this paper, we take the first steps of realizing such an object abstraction and explore storage mechanisms for these objects to enhance I/O performance, especially for scientific applications. We explore how an object-based interface can facilitate next generation scalable computing systems by presenting the mapping of data I/O from two real world HPC scientific use cases: a plasma physics simulation code (VPIC) and a cosmology simulation code (HACC). Our storage model stores data objects in different physical organizations to support data movement across layers of memory/storage hierarchy. Our implementation sclaes well to 16K parallel processes, and compared to the state of the art, such as MPI-IO and HDF5, our object-based data abstractions and data placement strategy in multi-level storage hierarchy achieves up to 7× I/O performance improvement for scientific data.