Managing Rich Metadata in High-Performance Computing Systems Using a Graph Model

Managing Rich Metadata in High-Performance Computing Systems Using a Graph Model
复制标题

DOI:
10.1109/tpds.2018.2887380
复制
发表时间:
2019-07
影响因子:
5.3
通讯作者:
Dong Dai;Yong Chen;P. Carns;John Jenkins;Wei Zhang;R. Ross
Dong Dai;Yong Chen;P. Carns;John Jenkins;Wei Zhang;R. Ross
中科院分区:
计算机科学2区
文献类型:
--
作者:
Dong Dai;Yong Chen;P. Carns;John Jenkins;Wei Zhang;R. Ross

文献摘要

被引文献

相似文献

高性能计算(HPC)系统生成大量关于不同实体(如作业、用户和文件)的元数据。现有系统可以有效地记录和管理这些元数据的一部分,主要是数据文件的POSIX元数据(如文件大小、名称和权限模式)。但是另一组重要的元数据,在本研究中被称为“富”元数据,它不仅记录了更广泛的实体(例如,运行的进程和作业),而且还记录了它们之间更复杂的关系,在当前的HPC系统中大多缺失。然而,这种丰富的元数据对于支持许多高级数据管理功能至关重要,例如识别给定结果背后的数据源和参数;审计数据使用情况;或者理解输入如何转化为输出的细节。为了统一有效地管理高性能计算系统中生成的丰富元数据,我们提出在本研究中使用图模型。我们确定了实现这样一个基于图的HPC富元数据管理系统的关键挑战,并提出了GraphMeta,一个为HPC平台设计和优化的基于图的富元数据管理系统,以解决这些挑战。对合成和真实HPC元数据工作负载的广泛评估表明,与现有解决方案相比,它在性能和可伸缩性方面都具有优势。
High-performance computing (HPC) systems generate huge amounts of metadata about different entities such as jobs, users, and files. Existing systems can efficiently record and manage part of these metadata, mainly the POSIX metadata of data files (e.g., file size, name, and permissions mode). But another important set of metadata, referred to as “rich” metadata in this study, which record not only wider range of entities (e.g., running processes and jobs) but also more complex relationships between them, are mostly missing in current HPC systems. Yet such rich metadata are critical for supporting many advanced data management functions such as identifying data sources and parameters behind a given result; auditing data usage; or understanding details about how inputs are transformed into outputs. To uniformly and efficiently manage the rich metadata generated in HPC systems, We propose to utilize a graph model in this study. We identify the key challenges of implementing such a graph-based HPC rich metadata management system and present GraphMeta, a graph-based rich metadata management system designed and optimized for HPC platforms, to tackle these challenges. Extensive evaluations on both synthetic and real HPC metadata workloads show its advantages in both performance and scalability compared with existing solutions.